bioRxiv ScienceSearch

Biology subjects

Hockenberry, A. J.

Publications and source records attributed to Hockenberry, A. J..

4 recordsLinked to original sources

Evolutionary couplings detect side-chain interactions

Patterns of amino acid covariation in large protein sequence alignments can inform the prediction of de novo protein structures, binding interfaces, and mutational effects. While algorithms that detect these so-called evolutionary couplings between residues have proven useful for practical applications, less is known about how and why these methods perform so well, and what insights into biological processes can be gained from their application. Evolutionary coupling algorithms are commonly benchmarked by comparison to true structural contacts derived from solved protein structures. However, the methods used to determine true structural contacts are not standardized and different definitions of structural contacts may have important consequences for interpreting the results from evolutionary coupling analyses and understanding their overall utility. Here, we show that evolutionary coupling analyses are significantly more likely to identify structural contacts between side-chain atoms than between backbone atoms. We use both simulations and empirical analyses to highlight that purely backbone-based definitions of true residue-residue contacts (i.e., based on the distance between C atoms) may underestimate the accuracy of evolutionary coupling algorithms by as much as 40% and that a commonly used reference point (C{beta} atoms) underestimates the accuracy by 10-15%. These findings show that co-evolutionary outcomes differ according to which atoms participate in residue-residue interactions and suggest that accounting for different interaction types may lead to further improvements to contact-prediction methods.\n\nSignificance StatementEvolutionary couplings between residues within a protein can provide valuable information about protein structures, protein-protein interactions, and the mutability of individual residues. However, the mechanistic factors that determine whether two residues will co-evolve remains unknown. We show that structural proximity by itself is not sufficient for co-evolution to occur between residues. Rather, evolutionary couplings between residues are specifically governed by interactions between side-chain atoms. By contrast, intramolecular contacts between atoms in the protein backbone display only a weak signature of evolutionary coupling. These findings highlight that different types of stabilizing contacts exist within protein structures and that these types have a differential impact on the evolution of protein structures that should be considered in co-evolutionary applications.

biophysics

Predicting bacterial growth conditions from mRNA and protein abundances

Cells respond to changing nutrient availability and external stresses by altering the expression of individual genes. Condition-specific gene expression patterns may provide a promising and low-cost route to quantifying the presence of various small molecules, toxins, or species-interactions in natural environments. However, whether gene expression signatures alone can predict individual environmental growth conditions remains an open question. Here, we used machine learning to predict 16 closely-related growth conditions using 155 datasets of E. coli transcript and protein abundances. We show that models are able to discriminate between different environmental features with a relatively high degree of accuracy. We observed a small but significant increase in model accuracy by combining transcriptome and proteome-level data, and we show that stationary phase conditions are typically more difficult to distinguish from one another than conditions under exponential growth. Nevertheless, with sufficient training data, gene expression measurements from a single species are capable of distinguishing between environmental conditions that are separated by a single environmental variable.

systems biology

Selection removes Shine-Dalgarno-like sequences from within protein coding genes

The Shine-Dalgarno (SD) sequence motif facilitates translation initiation and is frequently found upstream of bacterial start codons. However, thousands of instances of this motif occur throughout the middle of protein coding genes in a typical bacterial genome. Here, we use comparative evolutionary analysis to test whether SD sequences located within genes are functionally constrained. We measure the conservation of SD sequences across Gammaproteobacteria, and find that they are significantly less conserved than expected. Further, the strongest SD sequences are the least conserved whereas we find evidence of conservation for the weakest possible SD sequences given amino acid constraints. Our findings indicate that most SD sequences within genes are likely to be deleterious and removed via selection. To illustrate the origin of these deleterious costs, we show that ATG start codons are significantly depleted downstream of SD sequences within genes, highlighting the potential for these sequences to promote erroneous translation initiation.

evolutionary biology

Diversity of translation initiation mechanisms across bacterial species is driven by environmental conditions and growth demands

The Shine-Dalgarno (SD) sequence is often found upstream of protein coding genes across the bacterial kingdom, where it enhances start codon recognition via hybridization to the anti-SD (aSD) sequence on the small ribosomal subunit. Despite widespread conservation of the aSD sequence, the proportion of SD-led genes within a genome varies widely across species, and the evolutionary pressures shaping this variation remain largely unknown. Here, we conduct a phylogenetically-informed analysis and show that species capable of rapid growth have a significantly higher proportion of SD-led genes in their genome, suggesting a role for SD sequences in meeting the protein production demands of rapidly growing species. Further, we show that utilization of the SD sequence mechanism co-varies with: i) genomic traits that are indicative of efficient translation, and ii) optimal growth temperatures. In contrast to prior surveys, our results demonstrate that variation in translation initiation mechanisms across genomes is largely predictable, and that SD sequence utilization is part of a larger suite of translation-associated traits whose diversity is driven by the differential growth strategies of individual species.

systems biology