bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Evolutionary Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Phylodynamic assessment of intervention strategies for the West African Ebola virus outbreak

This preprint has been reviewed and recommended by Peer Community In Evolutionary Biology (http://dx.doi.org/10.24072/pci.evolbiol.100046). The recent Ebola virus (EBOV) outbreak in West Africa witnessed considerable efforts to obtain viral genomic data as the epidemic was unfolding. If such data can be deployed in real-time, molecular epidemiological investigations could play a role in complementing contact tracing undertaken by public health agencies. Analysing the EBOV genomes accumulated to date can also deliver insights into epidemic dynamics. Such analyses have been shown that metapopulation dynamics were critical for EBOV dispersal between rural and urban areas during the epidemic, but the implications for specific intervention scenarios remain unclear. Here, we address this issue using a collection of phylodynamic approaches. We show that long-distance dispersal events (between administrative areas >250 km apart) were not crucial for epidemic expansion and that preventing viral lineage movement to any given administrative area would, in most cases, have had little impact. However, urban areas - specifically those encompassing the three capital cities and their suburbs - were critical in attracting and further disseminating the virus: preventing viral lineage movement to all three simultaneously would have contained epidemic size by two-thirds. Using continuous phylogeographic reconstructions we estimate a distance kernel for EBOV spread and reveal considerable heterogeneity in dispersal velocity through time. We also show that announcements of border closures were followed by a significant but transient effect on international virus dispersal. By quantifying the hypothetical impact of different intervention strategies as well as the impact of barriers on dispersal frequency, our study illustrates how phylodynamic analyses can help to address specific epidemiological and outbreak control questions.

epidemiology

Error-prone bypass of DNA lesions during lagging strand replication is a common source of germline and cancer mutations

Spontaneously occurring mutations are of great relevance in diverse fields including biochemistry, oncology, evolutionary biology, and human genetics. Studies in experimental systems have identified a multitude of mutational mechanisms including DNA replication infidelity as well as many forms of DNA damage followed by inefficient repair or replicative bypass. However, the relative contributions of these mechanisms to human germline mutations remain completely unknown. Here, based on the mutational asymmetry with respect to the direction of replication and transcription, we suggest that error-prone damage bypass on the lagging strand plays a major role in human mutagenesis. Asymmetry with respect to transcription is believed to be mediated by the action of transcription-coupled DNA repair (TC-NER). TC-NER selectively repairs DNA lesions on the transcribed strand; as a result, lesions on the non-transcribed strand are preferentially converted into mutations. In human polymorphism we detect a striking similarity between transcriptional asymmetry and asymmetry with respect to replication fork direction. This parallels the observation that damage-induced mutations in human cancers accumulate asymmetrically with respect to the direction of replication, suggesting that DNA lesions are asymmetrically resolved during replication. Re-analysis of XR-seq data, Damage-seq data and cancers with defective NER corroborate the preferential error-prone bypass of DNA lesions on the lagging strand. We experimentally demonstrate that replication delay greatly attenuates the mutagenic effect of UV-irradiation, in line with the key role of replication in conversion of DNA damage to mutations. We conservatively estimate that at least 10% of human germline mutations arise due to DNA damage rather than replication infidelity. The number of these damage-induced mutations is expected to scale with the number of replications and, consequently, paternal age.

biochemistry

DiscoSnp-RAD: de novo detection of small variants for population genomics

We present an original method to de novo call variants for Restriction site associated DNA Sequencing (RAD-Seq). RAD-Seq is a technique characterized by the sequencing of specific loci along the genome, that is widely employed in the field of evolutionary biology since it allows to exploit variants (mainly SNPs) information from entire populations at a reduced cost. Common RAD dedicated tools, as STACKS or IPyRAD, are based on all-versus-all read comparisons, which require consequent time and computing resources. Based on the variant caller DiscoSnp, initially designed for shotgun sequencing, DiscoSnp-RAD avoids this pitfall as variants are detected by exploring the De Bruijn Graph built from all the read datasets. We tested the implementation on RAD data from 259 specimens of Chiastocheta flies, morphologically assigned to 7 species. All individuals were successfully assigned to their species using both STRUCTURE and Maximum Likelihood phylogenetic reconstruction. Moreover, identified variants succeeded to reveal a within species structuration and the existence of two populations linked to their geographic distributions. Furthermore, our results show that DiscoSnp-RAD is at least one order of magnitude faster than state-of-the-art tools. The overall results show that DiscoSnp-RAD is suitable to identify variants from RAD data, and stands out from other tools due to his completely different principle, making it significantly faster, in particular on large datasets.\n\nLicenseGNU Affero general public license\n\nAvailabilityhttps://github.com/GATB/DiscoSnp\n\nContactjeremy.gauthier@inria.fr

bioinformatics

Comparing phylogenetic trees according to tip label categories

Trees that illustrate patterns of ancestry and evolution are a central tool in many areas of biology. Comparing evolutionary trees to each other has widespread applications in comparing the evolutionary stories told by different sources of data, assessing the quality of inference methods, and highlighting areas where patterns of ancestry are uncertain. While these tasks are complicated by the fact that trees are high-dimensional structures encoding a large amount of information, there are a number of metrics suitable for comparing evolutionary trees whose tips have the same set of unique labels. There are also metrics for comparing trees where there is no relationship between their labels: in unlabelled tree metrics the tree shapes are compared without reference to the tip labels.\n\nIn many interesting applications, however, the taxa present in two or more trees are related but not identical, and it is informative to compare the trees whilst retaining information about their tips relationships. We present methods for comparing trees whose labels belong to a pre-defined set of categories. The methods include a measure of distance between two such trees, and a measure of concordance between one such tree and a hierarchical classification tree of the unique categories. We demonstrate the intuition of our methods with some toy examples before presenting an analysis of Mycobacterium tuberculosis trees, in which we use our methods to quantify the differences between trees built from typing versus sequence data.

evolutionary biology

Stochasticity of cellular growth: sources, propagation and consequences

Cellular growth impacts a range of phenotypic responses. Identifying the sources of fluctuations in growth and how they propagate across the cellular machinery can unravel mechanisms that underpin cell decisions. We present a stochastic cell model linking gene expression, metabolism and replication to predict growth dynamics in single bacterial cells. In addition to several population-averaged data, the model quantitatively recovers how growth fluctuations in single cells change across nutrient conditions. We develop a framework to analyse stochastic chemical reactions coupled with cell divisions and use it to identify sources of growth heterogeneity. By visualising cross-correlations we then determine how such initial fluctuations propagate to growth rate and affect other cell processes. We further study antibiotic responses and find that complex drug-nutrient interactions can both enhance and suppress heterogeneity. Our results provide a predictive framework to integrate single-cell and bulk data and draw testable predictions with implications for antibiotic tolerance, evolutionary biology and synthetic biology.

systems biology

Genie: An interactive real-time simulation for teaching genetic drift

Neutral evolution is a fundamental concept in evolutionary biology but teaching this and other non-adaptive concepts is specially challenging. Here we present Genie, a browser-based educational tool that facilitates demonstration of concepts such as genetic drift, population isolation, gene flow, and genetic mutation. Because it does not need to be downloaded and installed, Genie can scale to large groups of students and is useful for both in-person and online instruction. Genie was used to teach genetic drift to Evolution students at Arizona State University during Spring 2016 and Spring 2017. The effectiveness of Genie to teach key genetic drift concepts and misconceptions was assessed with the Genetic Drift Inventory developed by Price et al. (2014). Overall, Genie performed comparably to that of traditional static methods across all evaluated classes. We have empirically demonstrated that Genie can be successfully integrated with traditional instruction to reduce misconceptions about genetic drift.

scientific communication and education

Comparison of Genotypic and Phenotypic Correlations: Cheveruds Conjecture in Humans

Accurate estimation of genetic correlation requires large sample sizes and access to genetically informative data, which are not always available. Accordingly, phenotypic correlations are often assumed to reflect genotypic correlations in evolutionary biology. Cheveruds conjecture asserts that the use of phenotypic correlations as proxies for genetic correlations is appropriate. Empirical evidence of the conjecture has been found across plant and animal species, with results suggesting that there is indeed a robust relationship between the two. Here, we investigate the conjecture in human populations, an analysis made possible by recent developments in availability of human genomic data and computing resources. A sample of 108,035 British European individuals from the UK Biobank was split equally into discovery and replication datasets. 17 traits were selected based on sample size, distribution and heritability. Genetic correlations were calculated using linkage disequilibrium score regression applied to the genome-wide association summary statistics of pairs of traits, and compared within and across datasets. Strong and significant correlations were found for the between-dataset comparison, suggesting that the genetic correlations from one independent sample were able to predict the phenotypic correlations from another independent sample within the same population. Designating the selected traits as morphological or non-morphological indicated little difference in correlation. The results of this study support the existence of a relationship between genetic and phenotypic correlations in humans. This finding is of specific interest in anthropological studies, which use measured phenotypic correlations to make inferences about the genetics of ancient human populations.

genetics

A speed-fidelity trade-off determines the mutation rate and virulence of an RNA virus

Mutation rates can evolve through genetic drift, indirect selection due to genetic hitchhiking, or direct selection on the physicochemical cost of high fidelity. However, for many systems, it has been difficult to disentangle the relative impact of these forces empirically. In RNA viruses, an observed correlation between mutation rate and virulence has led many to argue that their extremely high mutation rates are advantageous, because they may allow for increased adaptability. This argument has profound implications, as it suggests that pathogenesis in many viral infections depends on rare or de novo mutations. Here we present data for an alternative model whereby RNA viruses evolve high mutation rates as a byproduct of selection for increased replicative speed. We find that a poliovirus antimutator, 3DG64S, has a significant replication defect and that wild type and 3DG64S populations have similar adaptability in two distinct cellular environments. Experimental evolution of 3DG64S under r-selection led to reversion and compensation of the fidelity phenotype. Mice infected with 3DG64S exhibited delayed morbidity at doses well above the LD50, consistent with attenuation by slower growth as opposed to reduced mutational supply. Furthermore, compensation of the 3DG64S growth defect restored virulence, while compensation of the fidelity phenotype did not. Our data are consistent with the kinetic proofreading model for biosynthetic reactions and suggest that speed is more important than accuracy. In contrast to what has been suggested for many RNA viruses, we find that within host spread is associated with viral replicative speed and not standing genetic diversity.\n\nAuthor SummaryMutation rate evolution has long been a fundamental problem in evolutionary biology. The polymerases of RNA viruses generally lack proofreading activity and exhibit extremely high mutation rates. Since most mutations are deleterious and mutation rates are tuned by natural selection, we asked why hasnt the virus evolved to have a lower mutation rate? We used experimental evolution and a murine infection model to show that RNA virus mutation rates may actually be too high and are not necessarily adaptive. Rather, our data indicate that viral mutation rates are driven higher as a result of selection for viruses with faster replication kinetics. We suggest that viruses have high mutation rates, not because they facilitate adaption, but because it is hard to be both fast and accurate.

microbiology

Methylation patterns reveal cryptic structure and a pathway for adaptation in a panmictic carnivore

Determining the molecular signatures of adaptive differentiation is a fundamental component of evolutionary biology. A key challenge remains for identifying such signatures in wild organisms, particularly between populations of highly mobile species that undergo substantial gene flow. The Canada lynx (Lynx canadensis) is one species where mainland populations appear largely undifferentiated at traditional genetic markers, despite inhabiting diverse environments and displaying phenotypic variation. Here, we used high-throughput sequencing to investigate both neutral genetic structure and epigenetic differentiation across the distributional range of Canada lynx. Using a customized bioinformatics pipeline we scored both neutral SNPs and methylated nucleotides across the lynx genome. Newfoundland lynx were identified as the most differentiated population at neutral genetic markers, with diffusion approximations of allele frequencies indicating that divergence from the panmictic mainland occurred at the end of the last glaciation, with minimal contemporary admixture. In contrast, epigenetic structure revealed hidden levels of differentiation across the range coincident with environmental determinants including winter conditions, particularly in the peripheral Newfoundland and Alaskan populations. Several biological pathways related to morphology were differentially methylated between populations, with Newfoundland being disproportionately methylated for genes that could explain the observed island dwarfism. Our results indicate that epigenetic modifications, specifically DNA methylation, are powerful markers to investigate population differentiation and functional plasticity in wild and non-model systems.\n\nSIGNIFICANCEPopulations experiencing high rates of gene flow often appear undifferentiated at neutral genetic markers, despite often extensive environmental and phenotypic variation. We examined genome-wide genetic differentiation and DNA methylation between three interconnected regions and one insular population of Canada lynx (Lynx canadensis) to determine if epigenetic modifications characterized climatic associations and functional molecular plasticity. Demographic approximations indicated divergence of Newfoundland during the last glaciation, while cryptic epigenetic structure identified putatively functional differentiation that might explain island dwarfism. Our study suggests that DNA methylation is a useful marker for differentiating wild populations, particularly when faced with functional plasticity and low genetic differentiation.

genomics

TreeSwift: a massively scalable Python package for trees

Phylogenetic trees are essential to evolutionary biology, and numerous methods exist that attempt to extract phylogenetic information applicable to a wide range of disciplines, such as epidemiology and metagenomics. Currently, the three main Python packages for trees are Bio.Phylo, DendroPy, and the ETE Toolkit, but as dataset sizes grow, parsing and manipulating ultra-large trees becomes impractical for these tools. To address this issue, we present TreeSwift, a user-friendly and massively scalable Python package for traversing and manipulating trees that is ideal for algorithms performed on ultra-large trees.

bioinformatics

Isolating and Quantifying the Role of Developmental Noise in Generating Phenotypic Variation

Phenotypic variation in organisms is typically attributed to genotypic variation, environmental variation, and their interaction. Developmental noise, which arises from stochasticity in cellular and molecular processes occurring during development when genotype and environment are fixed, also contributes to phenotypic variation. The potential influence of developmental noise is likely underestimated in studies of phenotypic variation due to intrinsic mechanisms within organisms that stabilize phenotypes and decrease variation. Since we are just beginning to appreciate the extent to which phenotypic variation due to stochasticity is potentially adaptive, the contribution of developmental noise to phenotypic variation must be separated and measured to fully understand its role in evolution. Here, we show that phenotypic variation due to genotype and environment, versus the contribution of developmental noise, can be distinguished for leopard gecko (Eublepharis macularius) head color patterns using mathematical simulations that model the role of random variation (corresponding to developmental noise) in patterning. Specifically, we modified the parameters of simulations corresponding to genetic and environmental variation to generate the full range of phenotypic variation in color pattern seen on the heads of eight leopard geckos. We observed that over the range of these parameters, the component of variation due to genotype and environment exceeds that due to developmental noise in the studied gecko cohort. However, the effect of developmental noise on patterning is also substantial. This approach can be applied to any regular morphological trait that results from self-organized processes such as reaction-diffusion mechanisms, including the frequently found striped and spotted patterns of animal pigmentation patterning, patterning of bones in vertebrate limbs, body segmentation in segmented animals. Our approach addresses one of the major goals of evolutionary biology: to define the role of stochasticity in shaping phenotypic variation.

developmental biology

Disentangling bacterial invasiveness from lethality in an experimental host-pathogen system

Quantifying virulence remains a central problem in human health, pest control, disease ecology, and evolutionary biology. Bacterial virulence is typically quantified by the LT50 (i.e. the time taken to kill 50% of infected hosts), however, such an indicator cannot account for the full complexity of the infection process, such as distinguishing between the pathogens ability to colonize vs. kill the hosts. Indeed, the pathogen needs to breach the primary defenses in order to colonize, find a suitable environment to replicate, and finally express the virulence factors that cause disease. Here, we show that two virulence attributes, namely pathogen lethality and invasiveness, can be disentangled from the survival curves of a laboratory population of Caenorhabditis elegans nematodes exposed to three bacterial pathogens: Pseudomonas aeruginosa, Serratia marcescens and Salmonella enterica. We first show that the host population eventually experiences a constant mortality rate, which quantifies the lethality of the pathogen. We then show that the time necessary to reach this constant-mortality rate regime depends on the pathogen growth rate and colonization rate, and thus determines the pathogen invasiveness. Our framework reveals that Serratia marcescens is particularly good at the initial colonization of the host, whereas Salmonella enterica is a poor colonizer yet just as lethal once established. Pseudomonas aeruginosa, on the other hand, is both a good colonizer and highly lethal after becoming established. The ability to quantitatively characterize the ability of different pathogens to perform each of these steps has implications for treatment and prevention of disease and for the evolution and ecology of pathogens.

ecology

GenoDup Pipeline: a tool to detect genome duplication using the dS-based method

Understanding whole genome duplication (WGD), or polyploidy, is fundamental to investigating the origin and diversification of organisms in evolutionary biology. The wealth of genomic data generated by next generation sequencing (NGS) has resulted in an urgent need for robust and accurate tools to detect WGD. Here, we present a useful and user-friendly pipeline called GenoDup for inferring WGD using the dS-based method. We have successfully applied GenoDup to identify WGD in empirical data from both plants and animals. The GenoDup Pipeline provides a reliable and useful tool to infer WGD from NGS data.

bioinformatics

Genome-wide signatures of local adaptation among seven stoneflies species along a nationwide latitudinal gradient in Japan

BackgroundEnvironmental heterogeneity continuously produces a selective pressure that results in genomic variation among organisms; understanding this relationship remains a challenge in evolutionary biology. Here, we evaluated the degree of genome-environmental association of seven stonefly species across a wide geographic area in Japan and additionally identified putative environmental drivers and their effect on co-existing multiple stonefly species. Double-digest restriction-associated DNA (ddRAD) libraries were independently sequenced for 219 individuals from 23 sites across four geographical regions along a nationwide latitudinal gradient in Japan.\n\nResultsA total of 4,251 candidate single nucleotide polymorphisms (SNPs) strongly associated with local adaptation were discovered using Latent mixed models; of these, 294 SNPs showed strong correlation with environmental variables, specifically precipitation and altitude, using distance-based redundancy analysis. Genome-genome comparison among the seven species revealed a high sequence similarity of candidate SNPs within a geographical region, suggesting the occurrence of a parallel evolution process.\n\nConclusionsOur results revealed genomic signatures of local adaptation and their influence on multiple, co-occurring species. These results can be potentially applied for future studies on river management and climatic stressor impacts.

genomics

The fate of deleterious variants in a barley genomic prediction population

Targeted identification and purging of deleterious genetic variants has been proposed as a novel approach to animal and plant breeding. This strategy is motivated, in part, by the observation that demographic events and strong selection associated with cultivated species pose a \"cost of domestication.\" This includes an increase in the proportion of genetic variants where a mutation is likely to reduce fitness. Recent advances in DNA resequencing and sequence constraint-based approaches to predict the functional impact of a mutation permit the identification of putatively deleterious SNPs (dSNPs) on a genome-wide scale. Using exome capture resequencing of 21 barley 6-row spring breeding lines, we identify 3,855 dSNPs among 497,754 total SNPs. In order to polarize SNPs as ancestral versus derived, we generated whole genome resequencing data of Hordeum murinum ssp. glaucum as a phylogenetic outgroup. The dSNPs occur at higher density in portions of the genome with a higher recombination rate than in pericentromeric regions with lower recombination rate and gene density. Using 5,215 progeny from a genomic prediction experiment, we examine the fate of dSNPs over three breeding cycles. Average derived allele frequency is lower for dSNPs than any other class of variants. Adjusting for initial frequency, derived alleles at dSNPs reduce in frequency or are lost more often than other classes of SNPs. The highest yielding lines in the experiment, as chosen by standard genomic prediction approaches, carry fewer homozygous dSNPs than randomly sampled lines from the same progeny cycle. In the final cycle of the experiment, progeny selected by genomic prediction have a mean of 5.6% fewer homozygous dSNPs relative to randomly chosen progeny from the same cycle.\n\nAuthor SummaryThe nature of genetic variants underlying complex trait variation has been the source of debate in evolutionary biology. Here, we provide evidence that agronomically important phenotypes are influenced by rare, putatively deleterious variants. We use exome capture resequencing and a hypothesis-based test for codon conservation to predict deleterious SNPs (dSNPS) in the parents of a multi-parent barley breeding population. We also generated whole-genome resequencing data of Hordeum murinum, a phylogenetic outgroup to barley, to polarize dSNPs by ancestral versus derived state. dSNPs occur disproportionately in the gene-rich chromosome arms, rather than in the recombination-poor pericentromeric regions. They also decrease in frequency more often than other variants at the same initial frequency during recurrent selection for grain yield and disease resistance. Finally, we identify a region on chromosome 4H that strongly associated with agronomic phenotypes in which dSNPs appear to be hitchhiking with favorable variants. Our results show that targeted identification and removal of dSNPs from breeding programs is a viable strategy for crop improvement, and that standard genomic prediction approaches may already contain some information about unobserved segregating dSNPs.

genomics

Comparison of evolutionary rescue via biological and cultural evolution

Rapid evolution allows populations to persist in environments where they would otherwise go extinct. This phenomenon, known as evolutionary rescue, is typically studied in the framework of biological evolution, yet adaptive traits can also arise and spread through cultural evolution. The present study developed a stochastic eco-evolutionary model to compare rescue probabilities through biological and cultural evolution. Transmission bias governed the rescue probability under cultural evolution by setting how readily a rare adaptive trait was copied. Conformity bias suppressed population persistence because a rare trait was the least likely to be copied. Content bias toward the adaptive trait enabled evolutionary rescue when social learning was rapid, but it typically yielded a lower rescue probability than biological evolution. Only anticonformity bias, together with a high social learning rate, exceeded the rescue probability of biological evolution by enabling the adaptive trait to be established more rapidly. These results demonstrate that transmission bias alters the demographic consequences of cultural evolution and highlight the importance of transmission processes in evolutionary rescue theory. Understanding how adaptive behaviours are socially transmitted may also improve predictions of animal population persistence and inform conservation efforts in rapidly changing environments.

evolutionary biology

Multi-omic analysis of a hyper-diverse plant metabolic pathway reveals evolutionary routes to biological innovation

The diversity of life on Earth is a result of continual innovations in molecular networks influencing morphology and physiology. Plant specialized metabolism produces hundreds of thousands of compounds, offering striking examples of these innovations. To understand how this novelty is generated, we investigated the evolution of the Solanaceae family-specific, trichome-localized acylsugar biosynthetic pathway using a combination of mass spectrometry, RNA-seq, enzyme assays, RNAi and phylogenetics in non-model species. Our results reveal that hundreds of acylsugars are produced across the Solanaceae family and even within a single plant, revealing this phenotype to be hyper-diverse. The relatively short biosynthetic pathway experienced repeated cycles of innovation over the last 100 million years that include gene duplication and divergence, gene loss, evolution of substrate preference and promiscuity. This study provides mechanistic insights into the emergence of plant chemical novelty, and offers a template for investigating the [~]300,000 non-model plant species that remain underexplored.

genomics

The necessary emergence of structural complexity in self-replicating RNA populations.

The RNA world hypothesis relies on the ability of ribonucleic acids to spontaneously acquire complex structures capable of supporting essential biological functions. Multiple sophisticated evolutionary models have been proposed for their emergence, but they often assume specific conditions. In this work we explore a simple and parsimonious scenario describing the emergence of complex molecular structures at the early stages of life. We show that at specific GC-content regimes, an undirected replication model is sufficient to explain the apparition of multi-branched RNA secondary structures - a structural signature of many essential ribozymes. We ran a large scale computational study to map energetically stable structures on complete mutational networks of 50-nucleotide-long RNA sequences. Our results reveal that the sequence landscape with stable structures is enriched with multi-branched structures at a length scale coinciding with the appearance of complex structures in RNA databases. A random replication mechanism preserving a 50% GC-content may suffice to explain a natural enrichment of stable complex structures in populations of functional RNAs. By contrast, an evolutionary mechanism eliciting the most stable folds at each generation appears to help reaching multi-branched structures at highest GC content.

evolutionary biology