bioRxiv ScienceSearch

Biology subjects

Beltrao, P.

Publications and source records attributed to Beltrao, P..

11 recordsLinked to original sources

Evolution of protein kinase substrate recognition at the active site

Protein kinases catalyse the phosphorylation of target proteins, controlling most cellular processes. The specificity of serine/threonine kinases is partly determined by interactions with a few residues near the phospho-acceptor residue, forming the so-called kinase substrate motif. Kinases have been extensively duplicated throughout evolution but little is known about when in time new target motifs have arisen. Here we show that sequence variation occurring early in the evolution of kinases is dominated by changes in specificity determining residues. We then analysed kinase specificity models, based on known target sites, observing that specificity has remained mostly unchanged for recent kinase duplications. Finally, analysis of phosphorylation data from a taxonomically broad set of 48 eukaryotic species indicates that most phosphorylation motifs are broadly distributed in eukaryotes but not present in prokaryotes. Overall, our results suggest that the set of eukaryotes kinase motifs present today was acquired soon after the eukaryotic last common ancestor and that early expansions of the protein kinase fold rapidly explored the space of possible target motifs.

evolutionary biology

Conserved phosphorylation hotspots in eukaryotic protein domain families

Protein phosphorylation is the best characterized post-translational modification that regulates almost all cellular processes through diverse mechanisms such as changing protein conformations, interactions, and localization. While the inventory for phosphorylation sites across different species has rapidly expanded, their functional role remains poorly investigated. Here, we have combined 537,321 phosphosites from 40 eukaryotic species to identify highly conserved phosphorylation \"hotspot\" regions within domain families. Mapping these regions onto structural data revealed that they are often found at interfaces, near catalytic residues and tend to harbor functionally important phosphosites. Notably, functional studies of a phospho-deficient mutant in the C-terminal hotspot region within the Ribosomal S11 domain in the yeast ribosomal protein uS11 showed cold-sensitive phenotype and impaired 20S pre-rRNA processing. Altogether, our study identified phosphorylation hotspots for 162 protein domains suggestive of an ancient role for the control of diverse eukaryotic domain families.

cell biology

iProteinDB: an integrative database of Drosophila post-translational modifications

Post-translational modification (PTM) serves as a regulatory mechanism for protein function, influencing stability, protein interactions, activity and localization, and is critical in many signaling pathways. The best characterized PTM is phosphorylation, whereby a phosphate is added to an acceptor residue, commonly serine, threonine and tyrosine. As proteins are often phosphorylated at multiple sites, identifying those sites that are important for function is a challenging problem. Considering that many phosphorylation sites may be non-functional, prioritizing evolutionarily conserved phosphosites provides a general strategy to identify the putative functional sites with regards to regulation and function. To facilitate the identification of conserved phosphosites, we generated a large-scale phosphoproteomics dataset from Drosophila embryos collected from six closely-related species. We built iProteinDB (https://www.flyrnai.org/tools/iproteindb/), a resource integrating these data with other high-throughput PTM datasets, including vertebrates, and manually curated information for Drosophila. At iProteinDB, scientists can view the PTM landscape for any Drosophila protein and identify predicted functional phosphosites based on a comparative analysis of data from closely-related Drosophila species. Further, iProteinDB enables comparison of PTM data from Drosophila to that of orthologous proteins from other model organisms, including human, mouse, rat, Xenopus laevis, Danio rerio, and Caenorhabditis elegans.

bioinformatics

Capturing variation impact on molecular interactions: the IMEx Consortium mutations data set

The current wealth of genomic variation data identified at the nucleotide level has provided us with the challenge of understanding by which mechanisms amino acid variation affects cellular processes. These effects may manifest as distinct phenotypic differences between individuals or result in the development of disease. Physical interactions between molecules are the linking steps underlying most, if not all, cellular processes. Understanding the effects that amino acid variation of a molecules sequence has on its molecular interactions is a key step towards connecting a full mechanistic characterization of nonsynonymous variation to cellular phenotype. Here we present an open access resource created by IMEx database curators over 14 years, featuring 28,000 annotations fully describing the effect of individual point sequence changes on physical protein interactions. We describe how this resource was built, the formats in which the data content is provided and offer a descriptive analysis of the data set. The data set is publicly available through the IntAct website at www.ebi.ac.uk/intact/resources/datasets#mutationDs and is being enhanced with every monthly release.

bioinformatics

Comprehensive variant effect predictions of single nucleotide variants in model organisms

The effect of single nucleotide variants (SNVs) in coding and non-coding regions is of great interest in genetics. Although many computational methods aim to elucidate the effects of SNVs on cellular mechanisms, it is not straightforward to comprehensively cover different molecular effects. To address this we compiled and benchmarked sequence and structure-based variant effect predictors and we analyzed the impact of nearly all possible amino acid and nucleotide variants in the reference genomes of H. sapiens, S. cerevisiae and E. coli. Studied mechanisms include protein stability, interaction interfaces, post-translational modifications and transcription factor binding sites. We apply this resource to the study of natural and disease coding variants. We also show how variant effects can be aggregated to generate protein complex burden scores that uncover protein complex to phenotype associations based on a set of newly generated growth profiles of 93 sequenced S. cerevisiae strains in 43 conditions. This resource is available through mutfunc, a tool by which users can query precomputed predictions by providing amino acid or nucleotide-level variants.

genomics

Global analysis of specificity determinants in eukaryotic protein kinases

Protein kinases lie at the heart of cell signalling processes, constitute one of the largest human domain families and are often mutated in disease. Kinase target recognition at the active site is in part determined by a few amino acids around the phosphoacceptor residue. These preferences vary across kinases and despite the increased knowledge of target substrates little is known about how most preferences are encoded in the kinase sequence and how these preferences evolve. Here, we used alignment-based approaches to identify 30 putative specificity determinant residues (SDRs) for 16 preferences. These were studied using structural models and were validated by activity assays of mutant kinases. Mutation data from patient cancer samples revealed that kinase specificity is often targeted in cancer to a greater extent than catalytic residues. Throughout evolution we observed that kinase specificity is strongly conserved across orthologs but can diverge after gene duplication as illustrated by the evolution of the G-protein coupled receptor kinase family. The identified SDRs can be used to predict kinase specificity from sequence and aid in the interpretation of evolutionary or disease-related genomic variants.

cell biology

Phenotype prediction in an Escherichia coli strain panel

Understanding how genetic variation contributes to phenotypic differences is a fundamental question in biology. Here, we set to predict fitness defects of an individual using mechanistic models of the impact of genetic variants combined with prior knowledge of gene function. We assembled a diverse panel of 696 Escherichia coli strains for which we obtained genomes and measured growth phenotypes in 214 conditions. We integrated variant effect predictors to derive gene-level probabilities of loss of function for every gene across strains. We combined these probabilities with information on conditional gene essentiality in the reference K-12 strain to predict the strains growth defects, providing significant predictions for up to 38% of tested conditions. The putative causal variants were validated in complementation assays highlighting commonly perturbed pathways in evolution for the emergence of growth phenotypes. Altogether, our work illustrates the power of integrating high-throughput gene function assays to predict the phenotypes of individuals.\n\nHighlightsO_LIAssembled a reference panel of E. coli strains\nC_LIO_LIGenotyped and high-throughput phenotyped the E. coli reference strain panel\nC_LIO_LIReliably predicted the impact of genetic variants in up to 38% of tested conditions\nC_LIO_LIHighlighted common genetic pathways for the emergence of deleterious phenotypes\nC_LI

systems biology

Sub-minute phosphoregulation of cell-cycle systems during Plasmodium gamete formation revealed by a high-resolution time course

Malaria parasites are protists of the genus Plasmodium, whose transmission to mosquitoes is initiated by the production of gametes. Male gametogenesis is an extremely rapid process that is tightly controlled to produce eight flagellated microgametes from a single haploid gametocyte within 10 minutes after ingestion by a mosquito. Regulation of the cell cycle is poorly understood in divergent eukaryotes like Plasmodium, where the highly synchronous response of gametocytes to defined chemical and physical stimuli from the mosquito has proved to be a powerful model to identify specific phosphorylation events critical for cell-cycle progression. To reveal the wider network of phosphorylation signalling in a systematic and unbiased manner, we have measured a high-resolution time course of the phosphoproteome of P. berghei gametocytes during the first minute of gametogenesis. The data show an extremely broad response in which distinct cell-cycle events such as initiation of DNA replication and mitosis are rapidly induced and simultaneously regulated. We identify several protein kinases and phosphatases that are likely central in the gametogenesis signalling pathway and validate our analysis by investigating the phosphoproteomes of mutants in two of them, CDPK4 and SRPK1. We show these protein kinases to have distinct influences over the phosphorylation of similar downstream targets that are consistent with their distinct cellular functions, which is revealed by a detailed phenotypic analysis of an SRPK1 mutant. Together, the results show that key cell-cycle systems in Plasmodium undergo simultaneous and rapid phosphoregulation. We demonstrate how a highly resolved time-course of dynamic phosphorylation events can generate deep insights into the unusual cell biology of a divergent eukaryote, which serves as a model for an important group of human pathogens.

microbiology

Chromosomal rearrangements are commonly post-transcriptionally attenuated in cancer

Chromosomal rearrangements, despite being detrimental, are ubiquitous in cancer and often act as driver events. The effect of copy number variations (CNVs) on the cellular proteome of tumours is poorly understood. Therefore, we have analysed recently generated proteogenomic data-sets on 282 tumour samples to investigate the impact of CNVs in the proteome of these cells. We found that CNVs are post-transcriptionally attenuated in 23-33% of proteins with an enrichment for protein complexes. Complex subunits are highly co-regulated and some act as rate-limiting steps of complex assembly, indirectly controlling the abundance of other complex members. We identified 48 such regulatory interactions and experimentally validated AP3B1 and GTF2E2 as controlling subunits. Lastly, we found that a gene-signature of protein attenuation is associated with increased resistance to chaperone and proteasome inhibitors. This study highlights the importance of post-transcriptional mechanisms in cancer which allow cells to cope with their altered genomes.

genomics

Genomic determinants of protein abundance variation in colorectal cancer cells

Assessing the extent to which genomic alterations compromise the integrity of the proteome is fundamental in identifying the mechanisms that shape cancer heterogeneity. We have used isobaric labelling and tribrid mass spectrometry to characterize the proteomic landscapes of 50 colorectal cancer cell lines and to decipher the relationships between genomic and proteomic variation. The robust quantification of 12,000 proteins and 27,000 phosphopeptides revealed how protein symbiosis translates to a co-variome which is subjected to a hierarchical order and exposes the collateral effects of somatic mutations on protein complexes. Targeted depletion of key chromatin modifiers confirmed the transmission of variation and the directionality as characteristics of protein interactions. Protein level variation was leveraged to build drug response predictive models towards a better understanding of pharmacoproteomic interactions in colorectal cancer. Overall, we provide a deep integrative view of the molecular structure underlying the variation of colorectal cancer cells.\n\nHighlightsO_LIThe cancer cell functional \"co-variome\" is a strong attribute of the proteome.\nC_LIO_LIMutations can have a direct impact on protein levels of chromatin modifiers.\nC_LIO_LITransmission of genomic variation is a characteristic of protein interactions.\nC_LIO_LIPharmacoproteomic models are strong predictors of response to DNA damaging agents.\nC_LI\n\nAbbreviations

systems biology

Benchmarking substrate-based kinase activity inference using phosphoproteomic data

MotivationPhosphoproteomic experiments are increasingly used to study the changes in signalling occurring across different conditions. It has been proposed that changes in phosphorylation of kinase target sites can be used to infer when a kinase activity is under regulation. However, these approaches have not yet been benchmarked due to a lack of appropriate benchmarking strategies.\n\nResultsWe curated public phosphoproteomic experiments to identify a gold standard dataset containing a total of 184 kinase-condition pairs where regulation is expected to occur. A list of kinase substrates was compiled and used to estimate changes in kinase activities using the following methods: Z-test, Kolmogorov Smirnov test, Wilcoxon rank sum test, gene set enrichment analysis (GSEA), and a multiple linear regression model (MLR). We also tested weighted variants of the Z-test, and GSEA that include information on kinase sequence specificity as proxy for affinity. Finally, we tested how the number of known substrates and the type of evidence (in vivo, in vitro or in silico) supporting these influence the predictions.\n\nConclusionsMost models performed well with the Z-test and the GSEA performing best as determined by the area under the ROc curve (Mean AUC=0.722). Weighting kinase targets by the kinase target sequence preference improves the results only marginally. However, the number of known substrates and the evidence supporting the interactions has a strong effect on the predictions.

bioinformatics