bioRxiv ScienceSearch

Biology subjects

Dmitri A Petrov

Publications and source records attributed to Dmitri A Petrov.

6 recordsLinked to original sources

Whole Genome Analysis of 132 Clinical Saccharomyces cerevisiae Strains Reveals Extensive Ploidy Variation

Budding yeast has undergone several independent transitions from commercial to clinical lifestyles. The frequency of such transitions suggests that clinical yeast strains are derived from environmentally available yeast populations, including commercial sources. However, despite their important role in adaptive evolution, the prevalence of polyploidy and aneuploidy has not extensively analyzed in clinical strains. In this study, we have looked for patterns governing the transition to clinical invasion in the largest screen of clinical yeast isolates to date. In particular, we have focused on the hypothesis that ploidy changes have influenced adaptive processes. We sequenced 145 yeast strains, 132 of which are clinical isolates. We found pervasive large-scale genomic variation in both overall ploidy (34% of strains identified as 3n/4n) and individual chromosomal copy numbers (36% of strains identified as aneuploid). We also found evidence for the highly dynamic nature of yeasts genomes, with 35 strains showing partial chromosomal copy number changes and 8 strains showing multiple independent chromosomal events. Intriguingly, a lineage identified to be baker/commercial derived with a unique damaging mutation in NDC80 was particularly prone to polyploidy, with 83% of its members being triploid or tetraploid. Polyploidy was in turn associated with a >2x increase in aneuploidy rates as compared to other lineages. This dataset provides a rich source of information of the genomics of clinical yeast strains and highlights the potential importance of large-scale genomic copy variation in yeast adaptation.

Genomics

Empirical evidence for heterozygote advantage in adapting diploid populations of Saccharomyces cerevisiae

Adaptation in diploids is predicted to proceed via mutations that are at least partially dominant in fitness. Recently we argued that many adaptive mutations might also be commonly overdominant in fitness. Natural (directional) selection acting on overdominant mutations should drive them into the population but then, instead of bringing them to fixation, should maintain them as balanced polymorphisms via heterozygote advantage. If true, this would make adaptive evolution in sexual diploids differ drastically from that of haploids. Unfortunately, the validity of this prediction has not yet been tested experimentally. Here we performed 4 replicate evolutionary experiments with diploid yeast populations (Saccharomyces cerevisiae) growing in glucose-limited continuous cultures. We sequenced 24 evolved clones and identified initial adaptive mutations in all four chemostats. The first adaptive mutations in all four chemostats were three CNVs, all of which proved to be overdominant in fitness. The fact that fitness overdominant mutations were always the first step in independent adaptive walks strongly supports the prediction that heterozygote advantage can arise as a common outcome of directional selection in diploids and demonstrates that overdominance of de novo adaptive mutations in diploids is not rare.

Evolutionary Biology

Viruses are a dominant driver of protein adaptation in mammals

Viruses interact with hundreds to thousands of proteins in mammals, yet adaptation against viruses has only been studied in a few proteins specialized in antiviral defense. Whether adaptation to viruses typically involves only specialized antiviral proteins or affects a broad array of proteins is unknown. Here, we analyze adaptation in ~1,300 virus-interacting proteins manually curated from a set of 9,900 proteins conserved across mammals. We show that viruses (i) use the more evolutionarily constrained proteins from the cellular functions they hijack and that (ii) despite this high constraint, virus-interacting proteins account for a high proportion of all protein adaptation in humans and other mammals. Adaptation is elevated in virus-interacting proteins across all functional categories, including both immune and non-immune functions. Our results demonstrate that viruses are one of the most dominant drivers of evolutionary change across mammalian and human proteomes.

Evolutionary Biology

More efficacious drugs lead to harder selective sweeps in the evolution of drug resistance in HIV-1

In the early days of HIV treatment, drug resistance occurred rapidly and predictably in all patients, but under modern treatments, resistance arises slowly, if at all. The probability of resistance should be controlled by the rate of generation of resistant mutations. If many adaptive mutations arise simultaneously, then adaptation proceeds by soft selective sweeps in which multiple adaptive mutations spread concomitantly, but if adaptive mutations occur rarely in the population, then a single adaptive mutation should spread alone in a hard selective sweep. Here we use 6,717 HIV-1 consensus sequences from patients treated with first-line therapies between 1989 and 2013 to confirm that the transition from fast to slow evolution of drug resistance was indeed accompanied with the expected transition from soft to hard selective sweeps. This suggests more generally that evolution proceeds via hard sweeps if resistance is unlikely and via soft sweeps if it is likely.

Evolutionary Biology

Illumina TruSeq synthetic long-reads empower de novo assembly and resolve complex, highly repetitive transposable elements

High-throughput DNA sequencing technologies have revolutionized genomic analysis, including the de novo assembly of whole genomes. Nevertheless, assembly of complex genomes remains challenging, in part due to the presence of dispersed repeats which introduce ambiguity during genome reconstruction. Transposable elements (TEs) can be particularly problematic, especially for TE families exhibiting high sequence identity, high copy number, or present in complex genomic arrangements. While TEs strongly affect genome function and evolution, most current de novo assembly approaches cannot resolve long, identical, and abundant families of TEs. Here, we applied a novel Illumina technology called TruSeq synthetic long-reads, which are generated through highly parallel library preparation and local assembly of short read data and achieve lengths of 1.5-18.5 Kbp with an extremely low error rate ([~]0.03% per base). To test the utility of this technology, we sequenced and assembled the genome of the model organism Drosophila melanogaster (reference genome strain y;cn,bw,sp) achieving an N50 contig size of 69.7 Kbp and covering 96.9% of the euchromatic chromosome arms of the current reference genome. TruSeq synthetic long-read technology enables placement of individual TE copies in their proper genomic locations as well as accurate reconstruction of TE sequences. We entirely recovered and accurately placed 4,229 (77.8%) of the 5,434 of annotated transposable elements with perfect identity to the current reference genome. As TEs are ubiquitous features of genomes of many species, TruSeq synthetic long-reads, and likely other methods that generate long reads, offer a powerful approach to improve de novo assemblies of whole genomes.

Genomics

Predictability of adaptive evolution under the successive fixation assumption

Predicting the course of evolution is critical for solving current biomedical challenges such as cancer and the evolution of drug resistant pathogens. One approach to studying evolutionary predictability is to observe repeated, independent evolutionary trajectories of similar organisms under similar selection pressures in order to empirically characterize this adaptive fitness landscape. As this approach is infeasible for many natural systems, a number of recent studies have attempted to gain insight into the adaptive fitness landscape by testing the plausibility of different orders of appearance for a specific set of adaptive mutations in a single adaptive trajectory. While this approach is technically feasible for systems with very few available adaptive mutations, the usefulness of this approach for predicting evolution in situations with highly polygenic adaptation is unknown. It is also unclear whether the presence of stable adaptive polymorphisms can influence the predictability of evolution as measured by these methods. In this work, we simulate adaptive evolution under Fishers geometric model to study evolutionary predictability. Remarkably, we find that the predictability estimated by these methods are anti-correlated, and that the presence of stable adaptive polymorphisms can both qualitatively and quantitatively change the predictability of evolution.

Evolutionary Biology