bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Evolutionary Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Estimation of sub-epidemic dynamics by means of Sequential Monte Carlo Approximate Bayesian Computation: an application to the Swiss HIV Cohort Study

Our ability to accurately infer transmission patterns of infectious diseases is critical to monitor both their spread and the efficacy of public health policies. The use of phylogenetic methods for the reconstruction of viral ancestral relationships has garnered increasing interest, particularly in the characterization of HIV epidemics and sub-epidemics. In the case of this virus, the Swiss HIV Cohort Study (SHCS) contains a wide breadth of genomic data that have been widely used as a means of applying such methods. However, current approaches for quantifying the epidemiological dynamics of diseases are computationally intensive, and fail to scale well with this magnitude of data. To address this issue, we re-implement an Approximate Bayesian Computation (ABC) approach based on sequential Monte Carlo (SMC). By means of simulations, we demonstrate that our implementation is capable of inferring key epidemiological parameters of the Swiss HIV epidemic accurately, and that sampling intensity has no significant effect on the accuracy of our estimates. Applied to a subset of HIV sequences from the SHCS, we show that we can distinguish sub-epidemics that are circulating in culturally distinct Swiss regions. Given these findings, we propose that ABC-SMC samplers will allow us to evaluate the impact of new public health policies, such as the implementation of a needle exchange program in the case of HIV, based on genetic data sampled before and after the implementation of a new policy.

evolutionary biology

Rapid and recent evolution of LTR retrotransposons drives rice genome evolution during the speciation of AA- genome Oryza species

The dynamics of LTR retrotransposons and their contribution to genome evolution during plant speciation have remained largely unanswered. Here, we perform a genome-wide comparison of all eight Oryza AA- genome species, and identify 3,911 intact LTR retrotransposons classified into 790 families. The top 44 most abundant LTR retrotransposon families show patterns of rapid and distinct diversification since the species split over the last ~4.8 Myr. Phylogenetic and read depth analyses of 11 representative retrotransposon families further provide a comprehensive evolutionary landscape of these changes. Compared with Ty1-copia, independent bursts of Ty3-gypsy retrotransposon expansions have occurred with the three largest showing signatures of lineage-specific evolution. The estimated insertion times of 2,213 complete retrotransposons from the top 23 most abundant families reveal divergent life-histories marked by speedy accumulation, decline and extinction that differed radically between species. We hypothesize that this rapid evolution of LTR retrotransposons not only divergently shaped the architecture of rice genomes but also contributed to the process of speciation and diversification of rice.

evolutionary biology

Direct estimation of the spontaneous mutation rate by short-term mutation accumulation lines in Chironomus riparius

Mutations are the ultimate basis of evolution, yet their occurrence rate is known only for few species. We directly estimated the spontaneous mutation rate and the mutational spectrum in the non-biting midge C. riparius with a new approach. Individuals from ten mutation accumulation lines over five generations were deep genome sequenced to count de novo mutations (DNMs) that were not present in a pool of F1 individuals, representing parental genotypes. We identified 51 new single site mutations of which 25 were insertions or deletions and 26 single point mutations. This shift in the mutational spectrum compared to other organisms was explained by the high A/T content of the species. We estimated a haploid mutation rate of 2.1 x 10-9 (95% confidence interval: 1.4 x 10-9 - 3.1 x 10-9) which is in the range of recent estimates for other insects and supports the drift barrier hypothesis. We show that accurate mutation rate estimation from a high number of observed mutations is feasible with moderate effort even for non-model species.

evolutionary biology

Buchnera has changed flatmate but the repeated replacement of co-obligate symbionts is not associated with the ecological expansions of their aphid hosts

Symbiotic associations with bacteria have facilitated important evolutionary transitions in insects and resulted in long-term obligate interactions. Recent evidence suggests that these associations are not always evolutionarily stable and that symbiont replacement and/or supplementation of an obligate symbiosis by an additional bacterium has occurred during the history of many insect groups. Yet, the factors favoring one symbiont over another in this evolutionary dynamic are not well understood; progress has been hindered by our incomplete understanding of the distribution of symbionts across phylogenetic and ecological contexts. While many aphids are engaged into an obligate symbiosis with a single Gammaproteobacterium, Buchnera aphidicola, in species of the Lachninae subfamily, this relationship has evolved into a \"menage a trois\", in which Buchnera is complemented by a cosymbiont, usually Serratia symbiotica. Using deep sequencing of 16S rRNA bacterial genes from 128 species of Cinara (the most diverse Lachninae genus), we reveal a highly dynamic dual symbiotic system in this aphid lineage. Most species host both Serratia and Buchnera but, in several clades, endosymbionts related to Sodalis, Erwinia or an unnamed member of the Enterobacteriaceae have replaced Serratia. Endosymbiont genome sequences from four aphid species+confirm that these coresident symbionts fulfill essential metabolic functions not ensured by Buchnera. We further demonstrate through comparative phylogenetic analyses that co-symbiont replacement is not associated with the adaptation of aphids to new ecological conditions. We propose that symbiont succession was driven by factors intrinsic to the phenomenon of endosymbiosis, such as rapid genome deterioration or competitive interactions between bacteria with similar metabolic capabilities.

evolutionary biology

Evolutionary forces affecting synonymous variations in plant genomes

Base composition is highly variable among and within plant genomes, especially at third codon positions, ranging from GC-poor and homogeneous species to GC-rich and highly heterogeneous ones (particularly Monocots). Consequently, synonymous codon usage is biased in most species, even when base composition is relatively homogeneous. The causes of these variations are still under debate, with three main forces being possibly involved: mutational bias, selection and GC-biased gene conversion (gBGC). So far, both selection and gBGC have been detected in some species but how their relative strength varies among and within species remains unclear. Population genetics approaches allow to jointly estimating the intensity of selection, gBGC and mutational bias. We extended a recently developed method and applied it to a large population genomic datasets based on transcriptome sequencing of 11 angiosperm species spread across the phylogeny. We found that base composition is far from mutation-drift equilibrium in most genomes and that gBGC is a widespread and stronger process than selection. gBGC could strongly contribute to base composition variation among plant species, implying that it should be taken into account in plant genome analyses, especially for GC-rich ones.

evolutionary biology

Anchored Phylogenomics of Angiosperms I: Assessing the Robustness of Phylogenetic Estimates

An important goal of the angiosperm systematics community has been to develop a shared approach to molecular data collection, such that phylogenomic data sets from different focal clades can be combined for meta-studies across the entire group. Although significant progress has been made through efforts such as DNA barcoding, transcriptome sequencing, and whole-plastid sequencing, the community current lacks a cost efficient methodology for collecting nuclear phylogenomic data across all angiosperms. Here, we leverage genomic resources from 43 angiosperm species to develop enrichment probes useful for collecting ~500 loci from non-model taxa across the diversity of angiosperms. By taking an anchored phylogenomics approach, in which probes are designed to represent sequence diversity across the group, we are able to efficiently target loci with sufficient phylogenetic signal to resolve deep, intermediate, and shallow angiosperm relationships. After demonstrating the utility of this resource, we present a method that generates a heat map for each node on a phylogeny that reveals the sensitivity of support for the node across analysis conditions, as well as different locus, site, and taxon schemes. Focusing on the effect of locus and site sampling, we use this approach to statistically evaluate relative support for the alternative relationships among eudicots, monocots, and magnoliids. Although the results from supermatrix and coalescent analyses are largely consistent across the tree, we find support for this deep relationship to be more sensitive to the particular choice of sites and loci when a supermatrix approach as employed. Averaged across analysis approaches and data subsampling schemes, our data support a eudicot-monocot sister relationship, which is supported by a number of recent angiosperm studies.

evolutionary biology

Biased gene conversion drives codon usage in human and precludes selection on translation efficiency

In humans, as in other mammals, synonymous codon usage (SCU) varies widely among genes. In particular, genes involved in cell differentiation or in proliferation display a distinct codon usage, suggesting that SCU is adaptively constrained to optimize translation efficiency in distinct cellular states. However, in mammals, SCU is known to correlate with large-scale fluctuations of GC-content along chromosomes, caused by meiotic recombination, via the non-adaptive process of GC-biased gene conversion (gBGC). To disentangle and to quantify the different factors driving SCU in humans, we analyzed the relationships between functional categories, base composition, recombination, and gene expression. We first demonstrate that SCU is predominantly driven by large-scale variation in GC-content and is not linked to constraints on tRNA abundance, which excludes an effect of translational selection. In agreement with the gBGC model, we show that differences in SCU among functional categories are explained by variation in intragenic recombination rate, which, in turn, is strongly negatively correlated to gene expression levels during meiosis. Our results indicate that variation in SCU among functional categories (including variation associated to differentiation or proliferation)- result from differences in levels of meiotic transcription, which interferes with the formation of crossovers and thereby affects gBGC intensity within genes. Overall, the gBGC model explains 81.3% of the variance in SCU among genes. We argue that the strong heterogeneity of SCU induced by gBGC in mammalian genomes precludes any optimization of the tRNA pool to the demand in codon usage.

evolutionary biology

Cladogenetic and Anagenetic Models of Chromosome Number Evolution: a Bayesian Model Averaging Approach

Chromosome number is a key feature of the higher-order organization of the genome, and changes in chromosome number play a fundamental role in evolution. Dysploid gains and losses in chromosome number, as well as polyploidization events, may drive reproductive isolation and lineage diversification. The recent development of probabilistic models of chromosome number evolution in the groundbreaking work by Mayrose et al. (2010, ChromEvol) have enabled the inference of ancestral chromosome numbers over molecular phylogenies and generated new interest in studying the role of chromosome changes in evolution. However, the ChromEvol approach assumes all changes occur anagenetically (along branches), and does not model events that are specifically cladogenetic. Cladogenetic changes may be expected if chromosome changes result in reproductive isolation. Here we present a new class of models of chromosome number evolution (called ChromoSSE) that incorporate both anagenetic and cladogenetic change. The ChromoSSE models allow us to determine the mode of chromosome number evolution; is chromosome evolution occurring primarily within lineages, primarily at lineage splitting, or in clade-specific combinations of both? Furthermore, we can estimate the location and timing of possible chromosome speciation events over the phylogeny. We implemented ChromoSSE in a Bayesian statistical framework, specifically in the software RevBayes, to accommodate uncertainty in parameter estimates while leveraging the full power of likelihood based methods. We tested ChromoSSEs accuracy with simulations and re-examined chromosomal evolution in Aristolochia, Carex section Spirostachyae, Helianthus, Mimulus sensu lato (s.l.), and Primula section Aleuritia, finding evidence for clade-specific combinations of anagenetic and cladogenetic dysploid and polyploid modes of chromosome evolution.

evolutionary biology

Revisiting the effect of red on competition in humans

Bright red coloration is a signal of male competitive ability in animal species across a range of taxa, including non-human primates. Does the effect of red on competition extend to humans? A landmark study in evolutionary psychology established such an effect through analysis of data for four combat sports at the 2004 Athens Olympics [1]. Here we show that the results do not replicate in an equivalent, independent dataset for the 2008 Beijing Olympics, and that there is substantial variation in the fraction of wins by red across sports in both years. We uncover a number of shortcomings with the research design, analysis, and interpretation underlying the original results. For example, the variation observed in the data may reflect bias towards wins by one color over the other, linked to specific features of the tournament structure for the sports analysed. Reanalysis of the data to address these shortcomings indicates that there is no evidence for an effect of red on the outcomes of Olympic combat sports. Our results refute past claims based on analysis of this system, challenging the related notion that any effect of red in human competition is an evolved response shaped by sexual selection.

evolutionary biology

Inactivation of thermogenic UCP1 as a historical contingency in multiple placental mammal clades

Mitochondrial uncoupling protein 1 (UCP1) is essential for non-shivering thermogenesis in brown adipose tissue and is widely accepted to have played a key thermoregulatory role in small-bodied and neonatal placental mammals that enabled the exploitation of cold environments. Here we map ucp1 sequences from 133 mammals onto a species tree constructed from a [~]51-kb sequence alignment and show that inactivating mutations have occurred in at least eight of the 18 traditional placental orders, thereby challenging the physiological importance of UCP1 across Placentalia. Selection and timetree analyses further reveal that ucp1 inactivations temporally correspond with strong secondary reductions in metabolic intensity in xenarthrans and pangolins, or in six other lineages coincided with a [~]30 million year episode of global cooling in the Paleogene that promoted sharp increases in body mass and cladogenesis evident in the fossil record. Our findings also demonstrate that members of various lineages (e.g., cetaceans, horses, woolly mammoths, Stellers sea cows) evolved extreme cold hardiness in the absence of UCP1-mediated thermogenesis. Finally, we identify ucp1 inactivation as a historical contingency that is linked to the current low species diversity of clades lacking functional UCP1, thus providing the first evidence for species selection related to the presence or absence of a single gene product.

evolutionary biology

Recurrent gene duplication leads to diverse repertoires of centromeric histones in Drosophila species

Despite their essential role in the process of chromosome segregation in most eukaryotes, centromeric histones show remarkable evolutionary lability. Not only have they been lost in multiple insect lineages, but they have also undergone gene duplication in multiple plant lineages. Based on detailed study of a handful of model organisms including Drosophila melanogaster, centromeric histone duplication is considered to be rare in animals. Using a detailed phylogenomic study, we find that Cid, the centromeric histone gene, has undergone four independent gene duplications during Drosophila evolution. We find duplicate Cid genes in D. eugracilis (Cid2), in the montium species subgroup (Cid3, Cid4) and in the entire Drosophila subgenus (Cid5). We show that Cid3, Cid4, Cid5 all localize to centromeres in their respective species. Some Cid duplicates are primarily expressed in the male germline. With rare exceptions, Cid duplicates have been strictly retained after birth, suggesting that they perform non-redundant centromeric functions, independent from the ancestral Cid. Indeed, each duplicate encodes a distinct N-terminal tail, which may provide the basis for distinct protein-protein interactions. Finally, we show some Cid duplicates evolve under positive selection whereas others do not. Taken together, our results support the hypothesis that Drosophila Cid duplicates have subfunctionalized. Thus, these gene duplications provide an unprecedented opportunity to dissect the multiple roles of centromeric histones.\n\nAuthor SummaryCentromeres ensure faithful segregation of DNA throughout eukaryotic life, thus providing the foundation for genetic inheritance. Paradoxically, centromeric proteins evolve rapidly despite being essential in many organisms. We have previously proposed that this rapid evolution is due to genetic conflict in female meiosis in which centromere alleles of varying strength compete for inclusion in the ovum. According to this centromere drive model, essential centromeric proteins (like the centromeric histone, CenH3) must evolve rapidly to counteract driving centromeres, which are associated with reduced male fertility. A simpler way to allow for the rapid evolution of centromeric proteins without compromising their essential function would be via gene duplication. Duplication and specialization of centromeric proteins would allow one paralog to function as a drive suppressor in the male germline, while allowing the other to carry out its canonical centromeric role. Here, we present the finding of multiple CenH3 (Cid) duplications in Drosophila. We identified four instances of Cid duplication followed by duplicate gene retention in Drosophila. These Cid duplicates were born between 20 and 40 million years ago. This finding more than doubles the number of known CenH3 duplications in animal species and suggests that most Drosophila species encode two or more Cid paralogs, in contrast to current view that most animal species only encode a single CenH3 gene. We show that duplicate Cid genes encode proteins that have retained the ability to localize to centromeres. We present three lines of evidence, which suggest that the multiple Cid duplications have been retained due to subfunctionalization. Based on these findings, we propose the novel hypothesis that the multiple functions carried out by CenH3 proteins, i.e., meiosis, mitosis and gametic inheritance, may be inherently incompatible with one another when encoded in a single locus.

evolutionary biology

An overview on the DNA nucleotide compositions across kingdoms

The DNA nucleotide compositions vary among species. This fascinating phenomenon has been studied for decades with some interesting questions remaining unclear. Recent years, thousands of genomes have been sequenced, but general evaluations on the nucleotide compositions across different phylogenetic groups are still absent. In this letter, I analyzed 371 genomes from different kingdoms and provided an overview on DNA nucleotide compositions. A number of important topics were discussed, including GC content, DNA strand symmetricity, CDS purine content, codon usage, thermophilicity in prokaryotes and non-coding RNA genes. I also gave explanations to two long debated questions: 1) both genome GC content and CDS purine content are correlated with the thermophilicity in archaea, but not in bacteria; 2) the purine rich pattern of CDS in most species is mainly a consequence of coding requirement, but not mRNA interaction dynamics. This study provides valuable information and ideas for future investigations.

evolutionary biology

The natural selection of metabolism and mass selects lifeforms from viruses to multicellular animals

I show that the natural selection of metabolism and mass is selecting for the major life history and allometric transitions that define lifeforms from viruses, over prokaryotes and larger unicells, to multicellular animals with sexual reproduction.\n\nThe proposed selection is driven by a mass specific metabolism that is selected as the pace of the resource handling that generates net energy for self-replication. This implies that an initial selection of mass is given by a dependence of mass specific metabolism on mass in replicators that are close to a lower size limit. A maximum dependence that is sublinear is shown to select for virus-like replicators with no intrinsic metabolism, no cell, and practically no mass. A maximum superlinear dependence is instead selecting for prokaryote-like self-replicating cells with asexual reproduction and incomplete metabolic pathways. These self-replicating cells have selection for increased net energy, and this generates a gradual unfolding of a population dynamic feed-back selection from interactive competition. The incomplete feed-back is shown to select for larger unicells with more developed metabolic pathways, and the completely developed feed-back to select for multicellular animals with sexual reproduction.\n\nThis model unifies natural selection from viruses to multicellular animals, and it provides a parsimonious explanation where allometries and major life history transitions evolve from the natural selection of metabolism and mass.

evolutionary biology

Functional constraints on replacing an essential gene with its ancient and modern homologs

The complexity hypothesis posits that network connectivity and protein function are two important determinants of how a gene adapts to and functions in a foreign genome. Genes encoding proteins that carry out essential informational tasks in the cell, in particular where multiple interaction partners are involved, are less likely to be transferable to a foreign organism. Here we investigated the constraints on transfer of a gene encoding a highly conserved informational protein, translation elongation factor Tu (EF-Tu), by systematically replacing the endogenous tufA gene in the Escherichia coli genome with its extant and ancestral homologs. The extant homologs represented tuf variants from both near and distant homologous organisms. The ancestral homologs represented phylogenetically resurrected tuf sequences dating from 0.7 to 3.6 bya. Our results demonstrate that all of the foreign tuf genes are transferable to the E. coli genome, provided that an additional copy of the EF-Tu gene, tufB, remains present in the E. coli genome. However, when the tufB gene was removed, only the variants obtained from the {gamma}-proteobacterial family (extant and ancestral), supported growth. This demonstrates the limited functional interchangability of E. coli tuf with its homologs. Our data show a linear correlation between relative bacterial fitness and the evolutionary distance of the extant tuf homologs inserted into the E. coli genome. Our data and analysis also suggest that the functional conservation of protein activity, and its network interactivity, act to constrain the successful transfer of this essential gene into foreign bacteria.

evolutionary biology

What can genomics tell us about the success of enhancement programs in anadromous Chinook salmon? A comparative analysis across four generations

Population enhancement through the release of cultured organisms can be an important tool for marine restoration. However, there has been considerable debate about whether releases effectively contribute to conservation and harvest objectives, and whether cultured organisms impact the fitness of wild populations. Pacific salmonid hatcheries on the West Coast of North America represent one of the largest enhancement programs in the world. Molecular-based pedigree studies on one or two generations have contributed to our understanding of the fitness of hatchery-reared individuals relative to wild individuals, and tend to show that hatchery fish have lower reproductive success. However, interpreting the significance of these results can be challenging because the long-term genetic and ecological effects of releases on supplemented populations are unknown. Further, pedigree studies have been opportunistic, rather than hypothesis driven, and have not provided information on \"best case\" management scenarios. Here, we present a comparative, experimental approach based on genome-wide surveys of changes in diversity in two hatchery lines founded from the same population. We demonstrate that gene flow with wild individuals can reduce divergence from the wild source population over four generations. We also report evidence for consistent genetic changes in a closed hatchery population that can be explained by both genetic drift and domestication selection. The results of this study suggest that genetic risks can be minimized over at least four generations with appropriate actions, and provide empirical support for a decision-making framework that is relevant to the management of hatchery populations.

evolutionary biology

Optimization of lag phase shapes the evolution of a bacterial enzyme

Mutations provide the variation that drives evolution, yet their effects on fitness remain poorly understood. Here we explore how mutations in the essential enzyme Adenylate Kinase (Adk) of E. coli affect multiple phases of population growth. We introduce a biophysical fitness landscape for these phases, showing how they depend on molecular and cellular properties of Adk. We find that Adk catalytic capacity in the cell (product of activity and abundance) is the major determinant of mutational fitness effects. We show that bacterial lag times are at a well-defined optimum with respect to Adks catalytic capacity, while exponential growth rates are only weakly affected by variation in Adk. Direct pairwise competitions between strains show how environmental conditions modulate the outcome of a competition where growth rates and lag times have a tradeoff, altogether shedding light on the multidimensional nature of fitness and its importance in the evolutionary optimization of enzymes.

evolutionary biology

Engineered reciprocal chromosome translocations drive high threshold, reversible population replacement in Drosophila

Replacement of wild insect populations with transgene-bearing individuals unable to transmit disease or survive under specific environmental conditions provides self-perpetuating methods of disease prevention and population suppression, respectively. Gene drive mechanisms that require the gene drive element and linked cargo exceed a high threshold frequency to spread are attractive because they offer several points of control: they bring about local, but not global population replacement; and transgenes can be eliminated by reintroducing wildtypes into the population so as to drive the frequency of transgenes below the threshold required for drive. It has long been recognized that reciprocal chromosome translocations could, in principal, be used to bring about high threshold gene drive through a form of underdominance. However, translocations able to drive population replacement have not been reported, leaving it unclear if translocation-bearing strains fit enough to mediate gene drive can easily be generated. Here we use modeling to identify a range of conditions under which translocations should spread, and the equilibrium frequencies achieved, given specific introduction frequencies, fitness costs and migration rates. We also report the creation of engineered translocation-bearing strains of Drosophila melanogaster, generated through targeted chromosomal breakage and homologous recombination. By several measures translocation-bearing strains are fit, and drive high threshold, reversible population replacement in laboratory populations. These observations, together with the generality of the tools used to generate translocations, suggest that engineered translocations may be useful for controlled population replacement in many species.

evolutionary biology

Overcoming evolved resistance to population-suppressing homing-based gene drives

The use of homing-based gene drive systems to modify or suppress wild populations of a given species has been proposed as a solution to a number of significant ecological and public health related problems, including the control of mosquito-borne diseases. The recent development of a CRISPR-Cas9-based homing system for the suppression of Anopheles gambiae, the main African malaria vector, is encouraging for this approach; however, with current designs, the slow emergence of homing-resistant alleles is expected to result in suppressed populations rapidly rebounding, as homing-resistant alleles have a significant fitness advantage over functional, population-suppressing homing alleles. To explore this concern, we develop a mathematical model to estimate tolerable rates of homing-resistant allele generation to suppress a wild population of a given size. Our results suggest that, to achieve meaningful population suppression, tolerable rates of resistance allele generation are orders of magnitude smaller than those observed for current designs for CRISPR-Cas9-based homing systems. To remedy this, we propose a homing system architecture in which guide RNAs (gRNAs) are multiplexed, increasing the effective homing rate and decreasing the effective resistant allele generation rate. Modeling results suggest that the size of the population that can be suppressed increases exponentially with the number of multiplexed gRNAs and that, with six multiplexed gRNAs, a mosquito species could potentially be suppressed on a continental scale. We also demonstrate successful multiplexing in vivo in Drosophila melanogaster using a ribozyme-gRNA-ribozyme (RGR) approach - a strategy that could readily be adapted to engineer stable, homing-based suppression drives in relevant organisms.

evolutionary biology