bioRxiv Science⌕ Search

Biology subjects

Dray, S.

Publications and source records attributed to Dray, S..

4 recordsLinked to original sources

PhylteR: efficient identification of outlier sequences in phylogenomic datasets

In phylogenomics, incongruences between gene trees, resulting from both artifactual and biological reasons, can decrease the signal-to-noise ratio and complicate species tree inference. The amount of data handled today in classical phylogenomic analyses precludes manual error detection and removal. However, a simple and efficient way to automate the identification of outliers from a collection of gene trees is still missing. Here, we present PhylteR, a method that allows a rapid and accurate detection of outlier sequences in phylogenomic datasets, i.e. species from individual gene trees that do not follow the general trend. PhylteR relies on DISTATIS, an extension of multidimensional scaling to 3 dimensions to compare multiple distance matrices at once. In PhylteR, these distance matrices extracted from individual gene phylogenies represent evolutionary distances between species according to each gene. On simulated datasets, we show that PhylteR identifies outliers with more sensitivity and precision than a comparable existing method. We also show that PhylteR is not sensitive to ILS-induced incongruences, which is a desirable feature. On a biological dataset of 14,463 genes for 53 species previously assembled for Carnivora phylogenomics, we show (i) that PhylteR identifies as outliers sequences that can be considered as such by other means, and (ii) that the removal of these sequences improves the concordance between the gene trees and the species tree. Thanks to the generation of numerous graphical outputs, PhylteR also allows for the rapid and easy visual characterisation of the dataset at hand, thus aiding in the precise identification of errors. PhylteR is distributed as an R package on CRAN and as containerized versions (docker and singularity).

evolutionary biology↗

Mycobacterium tuberculosis genetic features associated with pulmonary tuberculosis severity

Mycobacterium tuberculosis (Mtb) infections result in a wide spectrum of clinical presentations but without proven Mtb genetic determinants. Herein, 234 pulmonary tuberculosis (TB) patients were stratified according to TB disease severity and Mtb genetic features were explored using whole genome sequencing, including heterologous single nucleotide polymorphism (SNP) calling to explore micro-diversity. Clinical isolates from patients with mild TB carried mutations in genes associated with host-pathogen interaction, while those from patients with moderate/severe TB carried mutations associated with regulatory mechanisms. Genome-wide association study identified a SNP in the promoter of the gene coding for the virulence regulator EspR associated with moderate/severe disease. Structural equation modelling and model comparisons indicated that TB severity was associated with the detection of Mtb micro-diversity within clinical isolates and to the espR SNP. Taken together, these results provide a new insight to better understand TB pathophysiology and could provide new prognosis tool for pulmonary TB severity.

microbiology↗

Within-host genetic micro-diversity of Mycobacterium tuberculosis and the link with tuberculosis disease features

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb) complex, is still the number one deadly contagious disease. Mtb infection results in a wide spectrum of clinical presentations and severity symptoms, but without proven Mtb genetic determinants. Thanks to a collection of 355 clinical isolates with associated patients clinical data, we showed that Mtb micro-diversity within patient isolates is strongly correlated with TB-associated severity scores. Interestingly, this diversity is driven by a selection pressure to adapt to different lifestyles related to the infection site. Taken together, these results provide a new insight to better understand TB pathophysiology. Furthermore, Mtb micro-diversity could be envisioned as a new prognostic tool to improve the management of TB patients.

microbiology↗

Common causes drive negative correlations between nuclear genetic and species level biodiversity

The processes that give rise to species richness gradients are not well understood, but may be linked to resource-based limits on the number of species a region can support. Ecological limits placed on regional species richness would also limit population sizes, suggesting that these processes could also generate genetic diversity gradients. If true, we might better understand how broad-scale biodiversity patterns are formed by identifying the common causes of genetic diversity and species richness. We develop a hypothetical framework based on the consequences of regional variation in ecological limits to simultaneously explain spatial patterns of species richness and neutral genetic diversity. Repurposing raw genotypic data spanning 38 mammal species sampled across 801 sites in North America, we show that estimates of genome-wide genetic diversity and species richness share spatial structure. Notably, species richness hotspots tend to harbor lower levels of within-species genetic variation. A structural equation model encompassing eco-evolutionary processes related to resource availability, habitat heterogeneity, and human disturbance explained 78% of variation in genetic diversity and 74% of the variation in species richness. These results suggest we can infer broad-scale patterns of species and genetic diversity using two simple environmental measures of resource availability and ecological opportunity.

evolutionary biology↗