bioRxiv ScienceSearch

Biology subjects

Achaz, G.

Publications and source records attributed to Achaz, G..

5 recordsLinked to original sources

The genomic view of diversification

AO_SCPLOWBSTRACTC_SCPLOWEvolutionary relationships between species are traditionally represented in the form of a tree, the species tree. Its reconstruction from molecular data is hindered by frequent conflicts between gene genealogies. Usually, these disagreements are explained by incomplete lineage sorting (ILS) due to random coalescences of gene lineages inside the edges of the species tree. This paradigm, the multi-species coalescent (MSC), is constantly violated by the ubiquitous presence of gene flow, leading to incongruences between gene trees that cannot be explained by ILS alone. Here we argue instead in favor of a vision acknowledging the importance of gene flow and where gene histories shape the species tree rather than the opposite. We propose a new framework for modeling the joint evolution of gene and species lineages relaxing the hierarchy between the species tree and gene trees. We implement this framework in two mathematical models called the gene-based diversification models (GBD): 1) GBD-forward following all evolving genomes and 2) GBD-backward based on coalescent theory. They feature four parameters tuning colonization, gene flow, genetic drift and genetic differentiation. We propose a quick inference method based on differences between gene trees. Applied to two empirical data-sets prone to gene flow, we find a better support for the GBD model than for the MSC model. Along with the increasing awareness of the extent of gene flow, this work shows the importance of considering the richer signal contained in genomic histories, rather than in the mere species tree, to better apprehend the complex evolutionary history of species.

evolutionary biology

The quiescent X, the replicative Y and the Autosomes

From the analysis of the mutation spectrum in the 2,504 sequenced human genomes from the 1000 genomes project (phase 3), we show that sexual chromosomes (X and Y) exhibit a different proportion of indel mutations than autosomes (A), ranking them X>A>Y. We further show that X chromosomes exhibit a higher ratio of deletion/insertion when compared to autosomes. This simple pattern shows that the recent report that non-dividing quiescent yeast cells accumulate relatively more indels (and particularly deletions) than replicating ones also applies to metazoan cells, including humans. Indeed, the X chromosomes display more indels than the autosomes, having spent more time in quiescent oocytes, whereas the Y chromosomes are solely present in the replicating spermatocytes. From the proportion of indels, we have inferred that de novo mutations arising in the maternal lineage are twice more likely to be indels than mutations from the paternal lineage. Our observation, consistent with a recent trio analysis of the spectrum of mutations inherited from the maternal lineage, is likely a major component in our understanding of the origin of anisogamy.

evolutionary biology

Quiescence unveils a novel mutational force in fission yeast

One Sentence SummaryThe quiescence-driven mutational landscape reveals a novel evolutionary force.\n\nAbstractDuring cell division, the spontaneous mutation rate is expressed as the probability of mutations per generation, whereas during quiescence it will be expressed per unit of time. In this study, we report that during quiescence, the unicellular haploid fission yeast accumulates mutations as a linear function of time. We determined that 3 days of quiescence generate a number of invalidating mutations equivalent to that of one round of DNA replication. The novel mutational landscape of quiescence is characterized by insertion/deletion accumulating as fast as single nucleotide variants, and elevated amounts of deletions. When we extended the study to 3 months of quiescence, we confirmed the replication-independent mutational spectrum at the whole-genome level of a clonally aged population and uncovered phenotypic variations that subject the cells to natural selection. Thus, our results support the idea that genomes continuously evolve under two alternating phases that will impact on their size and composition.

genomics

A Molecular Portrait Of De Novo Genes

New genes, with novel protein functions, can evolve \"from scratch\" out of intergenic sequences. These de novo genes can integrate the cells genetic network and drive important phenotypic innovations. Therefore, identifying de novo genes and understanding how the transition from noncoding to coding occurs are key problems in evolutionary biology. However, identifying de novo genes is a difficult task, hampered by the presence of remote homologs, fast evolving sequences and erroneously annotated protein coding genes. To overcome these limitations, we developed a procedure that handles the usual pitfalls in de novo gene identification and predicted the emergence of 703 de novo genes in 15 yeast species from two genera whose phylogeny spans at least 100 million years of evolution. We established that de novo gene origination is a widespread phenomenon in yeasts, only a few being ultimately maintained by selection. We validated 82 candidates, by providing new translation evidence for 25 of them through mass spectrometry experiments. We also unambiguously identified the mutations that enabled the transition from non-coding to coding for 30 Saccharomyces de novo genes. We found that de novo genes preferentially emerge next to divergent promoters in GC-rich intergenic regions where the probability of finding a fortuitous and transcribed ORF is the highest. We found a more than 3-fold enrichment of de novo genes at recombination hot spots, which are GC-rich and nucleosome-free regions, suggesting that meiotic recombination would be a major driving force of de novo gene emergence in yeasts.

evolutionary biology

The expected neutral frequency spectrum of linked sites

We introduce the conditional Site Frequency Spectrum (SFS) for a genomic region linked to a focal mutation of known frequency. An exact expression for its expected value is provided for the neutral model without recombination. Its relation with the expected SFS for two sites, 2-SFS, is discussed. These spectra derive from the coalescent approach of Fu (1995) for finite samples, which is reviewed. Remarkably simple expressions are obtained for the linked SFS of a large population, which are also solutions of the multiallelic Kolmogorov equations. These formulae are the immediate extensions of the well known single site{theta} /f neutral SFS. Besides the general interest in these spectra, they relate to relevant biological cases, such as structural variants and introgressions. As an application, a recipe to adapt Tajimas D and other SFS-based neutrality tests to a non-recombining region containing a neutral marker is presented.

genetics