bioRxiv ScienceSearch

Biology subjects

Sarah, G.

Publications and source records attributed to Sarah, G..

3 recordsLinked to original sources

Pervasive hybridizations in the history of wheat relatives

Bread wheat and durum wheat derive from an intricate evolutionary history of three genomes, namely A, B and D, present in both extent diploid and polyploid species. Despite its importance for wheat research, no consensus on the phylogeny of the wheat clade has emerged so far, possibly because of hybridizations and gene flows that make phylogeny reconstruction challenging. Recently, it has been proposed that the D genome originated from an ancient hybridization event between the A and B genomes1. However, the study only relied on four diploid wheat relatives when 13 species are accessible. Using transcriptome data from all diploid species and a new methodological approach, we provide the first comprehensive phylogenomic analysis of this group. Our analysis reveals that most species belong to the D-genome lineage and descend from the previously detected hybridization event, but with a more complex scenario and with a different parent than previously thought. If we confirmed that one parent was the A genome, we found that the second was not the B genome but the ancestor of Aegilops mutica (T genome), an overlooked wild species. We also unravel evidence of other massive gene flow events that could explain long-standing controversies in the classification of wheat relatives. We anticipate that these results will strongly affect future wheat research by providing a robust evolutionary framework and refocusing interest on understudied species. The new method we proposed should also be pivotal for further methodological developments to reconstruct species relationship with multiple hybridizations.

evolutionary biology

TOGGLe, a flexible framework for easily building complex workflows and performing robust large-scale NGS analyses

The advent of NGS has intensified the need for robust pipelines to perform high-performance automated analyses. The required softwares depend on the sequencing method used to produce raw data (e.g. Whole genome sequencing, Genotyping By Sequencing, RNASeq) as well as the kind of analyses to carry on (GWAS, population structure, differential expression). These tools have to be generic and scalable, and should meet the biologists needs.\n\nHere, we present the new version of TOGGLe (Toolbox for Generic NGS Analyses), a simple and highly flexible framework to easily and quickly generate pipelines for large-scale second-and third-generation sequencing analyses, including multi-threading support. TOGGLe comprises a workflow manager designed to be as effortless as possible to use for biologists, so the focus can remain on the analyses. Embedded pipelines are easily customizable and supported analyses are reproducible and shareable. TOGGLe is designed as a generic, adaptable and fast evolutive solution, and has been tested and used in large-scale projects with numerous samples and organisms. It is freely available at http://toggle.southgreen.fr/ under the GNU GPLv3/CeCill-C licenses) and can be deployed onto HPC clusters as well as on local machines.

bioinformatics

Evolutionary forces affecting synonymous variations in plant genomes

Base composition is highly variable among and within plant genomes, especially at third codon positions, ranging from GC-poor and homogeneous species to GC-rich and highly heterogeneous ones (particularly Monocots). Consequently, synonymous codon usage is biased in most species, even when base composition is relatively homogeneous. The causes of these variations are still under debate, with three main forces being possibly involved: mutational bias, selection and GC-biased gene conversion (gBGC). So far, both selection and gBGC have been detected in some species but how their relative strength varies among and within species remains unclear. Population genetics approaches allow to jointly estimating the intensity of selection, gBGC and mutational bias. We extended a recently developed method and applied it to a large population genomic datasets based on transcriptome sequencing of 11 angiosperm species spread across the phylogeny. We found that base composition is far from mutation-drift equilibrium in most genomes and that gBGC is a widespread and stronger process than selection. gBGC could strongly contribute to base composition variation among plant species, implying that it should be taken into account in plant genome analyses, especially for GC-rich ones.

evolutionary biology