bioRxiv ScienceSearch

Biology subjects

Ochoa, D.

Publications and source records attributed to Ochoa, D..

3 recordsLinked to original sources

Conserved phosphorylation hotspots in eukaryotic protein domain families

Protein phosphorylation is the best characterized post-translational modification that regulates almost all cellular processes through diverse mechanisms such as changing protein conformations, interactions, and localization. While the inventory for phosphorylation sites across different species has rapidly expanded, their functional role remains poorly investigated. Here, we have combined 537,321 phosphosites from 40 eukaryotic species to identify highly conserved phosphorylation \"hotspot\" regions within domain families. Mapping these regions onto structural data revealed that they are often found at interfaces, near catalytic residues and tend to harbor functionally important phosphosites. Notably, functional studies of a phospho-deficient mutant in the C-terminal hotspot region within the Ribosomal S11 domain in the yeast ribosomal protein uS11 showed cold-sensitive phenotype and impaired 20S pre-rRNA processing. Altogether, our study identified phosphorylation hotspots for 162 protein domains suggestive of an ancient role for the control of diverse eukaryotic domain families.

cell biology

Capturing variation impact on molecular interactions: the IMEx Consortium mutations data set

The current wealth of genomic variation data identified at the nucleotide level has provided us with the challenge of understanding by which mechanisms amino acid variation affects cellular processes. These effects may manifest as distinct phenotypic differences between individuals or result in the development of disease. Physical interactions between molecules are the linking steps underlying most, if not all, cellular processes. Understanding the effects that amino acid variation of a molecules sequence has on its molecular interactions is a key step towards connecting a full mechanistic characterization of nonsynonymous variation to cellular phenotype. Here we present an open access resource created by IMEx database curators over 14 years, featuring 28,000 annotations fully describing the effect of individual point sequence changes on physical protein interactions. We describe how this resource was built, the formats in which the data content is provided and offer a descriptive analysis of the data set. The data set is publicly available through the IntAct website at www.ebi.ac.uk/intact/resources/datasets#mutationDs and is being enhanced with every monthly release.

bioinformatics

Benchmarking substrate-based kinase activity inference using phosphoproteomic data

MotivationPhosphoproteomic experiments are increasingly used to study the changes in signalling occurring across different conditions. It has been proposed that changes in phosphorylation of kinase target sites can be used to infer when a kinase activity is under regulation. However, these approaches have not yet been benchmarked due to a lack of appropriate benchmarking strategies.\n\nResultsWe curated public phosphoproteomic experiments to identify a gold standard dataset containing a total of 184 kinase-condition pairs where regulation is expected to occur. A list of kinase substrates was compiled and used to estimate changes in kinase activities using the following methods: Z-test, Kolmogorov Smirnov test, Wilcoxon rank sum test, gene set enrichment analysis (GSEA), and a multiple linear regression model (MLR). We also tested weighted variants of the Z-test, and GSEA that include information on kinase sequence specificity as proxy for affinity. Finally, we tested how the number of known substrates and the type of evidence (in vivo, in vitro or in silico) supporting these influence the predictions.\n\nConclusionsMost models performed well with the Z-test and the GSEA performing best as determined by the area under the ROc curve (Mean AUC=0.722). Weighting kinase targets by the kinase target sequence preference improves the results only marginally. However, the number of known substrates and the evidence supporting the interactions has a strong effect on the predictions.

bioinformatics