bioRxiv ScienceSearch

Biology subjects

Marks, D. S.

Publications and source records attributed to Marks, D. S..

8 recordsLinked to original sources

The EVcouplings Python framework for coevolutionary sequence analysis

SummaryCoevolutionary sequence analysis has become a commonly used technique for de novo prediction of the structure and function of proteins, RNA, and protein complexes. This approach requires extensive computational pipelines that integrate multiple tools, databases, and data processing steps. We present the EVcouplings framework, a fully integrated open-source application and Python package for coevolutionary analysis. The framework enables generation of sequence alignments, calculation and evaluation of evolutionary couplings (ECs), and de novo prediction of structure and mutation effects. The application has an easy to use command line interface to run workflows with user control over all analysis parameters, while the underlying modular Python package allows interactive data analysis and rapid development of new workflows. Through this multi-layered approach, the EVcouplings framework makes the full power of coevolutionary analyses available to entry-level and advanced users.\n\nAvailabilityhttps://github.com/debbiemarkslab/evcouplings\n\nContactsander.research@gmail.com, debbie@hms.harvard.edu

bioinformatics

Genome-wide discovery of epistatic loci affecting antibiotic resistance using evolutionary couplings

The analysis of whole genome sequencing data should, in theory, allow the discovery of interdependent loci that cause antibiotic resistance. In practice, however, identifying this epistasis remains a challenge as the vast number of possible interactions erodes statistical power. To solve this problem, we extend a method that has been successfully used to identify epistatic residues in proteins to infer genomic loci that are strongly coupled and associated with antibiotic resistance. Our method reduces the number of tests required for an epistatic genome-wide association study and increases the likelihood of identifying causal epistasis. We discovered 38 loci and 250 epistatic pairs that influence the dose needed to inhibit growth for five different antibiotics in 1,102 isolates of Neisseria gonorrhoeae that were confirmed in an independent dataset of 495 isolates. Many known resistance-affecting loci were recovered; however, the majority of loci occurred in unreported genes, including murE which was associated with cefixime. About half of the novel epistasis we report involved at least one locus previously associated with antibiotic resistance, including interactions between gyrA and parC associated with ciprofloxacin. Still, many combinations involved unreported loci and genes. Our work provides a systematic identification of epistasis pairs affecting antibiotic resistance in N. gonorrhoeae and a generalizable method for epistatic genome-wide association studies.

genomics

3D protein structure from genetic epistasis experiments

High-throughput experimental techniques have made possible the systematic sampling of the single mutation landscape for many proteins, defined as the change in protein fitness as the result of point mutation sequence changes. In a more limited number of cases, and for small proteins only, we also have nearly full coverage of all possible double mutants. By comparing the phenotypic effect of two simultaneous mutations with that of the individual amino acid changes, we can evaluate epistatic effects that reflect non-additive cooperative processes. The observation that epistatic residue pairs often are in contact in the 3D structure led to the hypothesis that a systematic epistatic screen contains sufficient information to identify the 3D fold of a protein. To test this hypothesis, we examined experimental double mutants for evidence of epistasis and identified residue contacts at 86% accuracy, including secondary structure elements and evidence for an alternative all--helical conformation. Positively epistatic contacts - corresponding to compensatory mutations, restoring fitness - were the most informative. Folded models generated from top-ranked epistatic pairs, when compared with the known structure, were accurate within 2.4 [A] over 53 residues, indicating the possibility that 3D protein folds can be determined experimentally with good accuracy from functional assays of mutant libraries, at least for small proteins. These results suggest a new experimental approach for determining protein structure.

systems biology

Structure and mutagenic analysis of the lipid II flippase MurJ from Escherichia coli

The peptidoglycan cell wall provides an essential protective barrier in almost all bacteria, defining cellular morphology and conferring resistance to osmotic stress and other environmental hazards. The precursor to peptidoglycan, lipid II, is assembled on the inner leaflet of the plasma membrane. However, peptidoglycan polymerization occurs on the outer face of the plasma membrane, and lipid II must be flipped across the membrane by the MurJ protein prior to its use in peptidoglycan synthesis. Due to its central role in cell wall assembly, MurJ is of fundamental importance in microbial cell biology and is a prime target for novel antibiotic development. However, relatively little is known regarding the mechanisms of MurJ function, and structural data are only available for MurJ from the extremophile Thermosipho africanus. Here, we report the crystal structure of substrate-free MurJ from the Gram-negative model organism Escherichia coli, revealing an inward-open conformation. Taking advantage of the genetic tractability of E. coli, we performed high-throughput mutagenesis and next-generation sequencing to assess mutational tolerance at every amino acid in the protein, providing a detailed functional and structural map for the enzyme and identifying sites for inhibitor development. Finally, through the use of sequence co-evolution analysis we identify functionally important interactions in the outward-open state of the protein, supporting a rocker-switch model for lipid II transport.

biochemistry

Viral gain-of-function experiments uncover residues under diversifying selection in nature

Viral gain-of-function mutations are commonly observed in the laboratory; however, it is unknown whether those mutations also evolve in nature. We identify two key residues in the host recognition protein of bacteriophage {lambda} that are necessary to exploit a new receptor; both residues repeatedly evolved among homologs from environmental samples. Our results provide evidence for widespread host-shift evolution in nature and a proof of concept for integrating experiments with genomic epidemiology.

evolutionary biology

Deep generative models of genetic variation capture mutation effects

The functions of proteins and RNAs are determined by a myriad of interactions between their constituent residues, but most quantitative models of how molecular phenotype depends on genotype must approximate this by simple additive effects. While recent models have relaxed this constraint to also account for pairwise interactions, these approaches do not provide a tractable path towards modeling higher-order dependencies. Here, we show how latent variable models with nonlinear dependencies can be applied to capture beyond-pairwise constraints in biomolecules. We present a new probabilistic model for sequence families, DeepSequence, that can predict the effects of mutations across a variety of deep mutational scanning experiments significantly better than site independent or pairwise models that are based on the same evolutionary data. The model, learned in an unsupervised manner solely from sequence information, is grounded with biologically motivated priors, reveals latent organization of sequence families, and can be used to extrapolate to new parts of sequence space.

genetics

Noise control is a primary function of microRNAs and post-transcriptional regulation

microRNAs are pervasive post-transcriptional regulators of protein-coding genes in multicellular organisms. Two fundamentally different models have been proposed for the function of microRNAs in gene regulation. In the first model, microRNAs act as repressors, reducing protein concentrations by accelerating mRNA decay and inhibiting translation. In the second model, in contrast, the role of microRNAs is not to reduce protein concentrations per se but to reduce fluctuations in these concentrations. Here we present genome-wide evidence that mammalian microRNAs frequently function as noise controllers rather than repressors. Moreover, we show that post-transcriptional noise control has been widely adopted across species from bacteria to animals, with microRNAs specifically employed to reduce noise in regulatory and context-specific processes in animals. Our results substantiate the detrimental nature of expression noise, reveal a universal strategy to control it, and suggest that microRNAs represent an evolutionary innovation for adaptive noise control in animals.\n\nHighlightsO_LIGenome-wide evidence that microRNAs function as noise controllers for genes with context-specific functions\nC_LIO_LIPost-transcriptional noise control is universal from bacteria to animals\nC_LIO_LIAnimals have evolved noise control for regulatory and context-specific processes\nC_LI

genetics

Genetic variation in human drug-related genes

Variability in drug efficacy and adverse effects are observed in clinical practice. While the extent of genetic variability in classical pharmacokinetic genes is rather well understood, the role of genetic variation in drug targets is typically less studied. Based on 60,706 human exomes from the ExAC dataset, we performed an in-depth computational analysis of the prevalence of functional-variants in in 806 drug-related genes, including 628 known drug targets. We find that most genetic variants in these genes are very rare (f < 0.1%) and thus likely not observed in clinical trials. Overall, however, four in five patients are likely to carry a functional-variant in a target for commonly prescribed drugs and many of these might alter drug efficacy. We further computed the likelihood of 1,236 FDA approved drugs to be affected by functional-variants in their targets and show that the patient-risk varies for many drugs with respect to geographic ancestry. A focused analysis of oncological drug targets indicates that the probability of a patient carrying germline variants in oncological drug targets is with 44% high enough to suggest that not only somatic alterations, but also germline variants carried over into the tumor genome should be included in therapeutic decision-making.

bioinformatics