bioRxiv ScienceSearch

Biology subjects

Aris-Brosou, S.

Publications and source records attributed to Aris-Brosou, S..

5 recordsLinked to original sources

Evidence of a nonadaptive buildup of mutational load in human populations over the past 40,000 years

The extent to which selection has shaped present-day human populations has attracted intense scrutiny, and examples of local adaptations abound. However, the evolutionary trajectory of alleles that, today, are deleterious has received much less attention. To address this question, the genomes of 2,062 individuals, including 1,179 ancient humans, were reanalyzed to assess how frequencies of risk alleles and their homozygosity changed through space and time in Europe over the past 45,000 years. While the overall deleterious homozygosity has consistently decreased, risk alleles have steadily increased in frequency over that period of time. Those that increased most are associated with diseases such as asthma, Crohn disease, diabetes and obesity, which are highly prevalent in present-day populations. These findings may not run against the existence of local adaptations, but highlight the limitations imposed by drift and population dynamics on the strength of selection in purging deleterious mutations from human populations.

evolutionary biology

How the Central American Seaway and an ancient northern passage affected Flatfish diversification

While the natural history of flatfish has been debated for decades, the mode of diversification of this biologically and economically important group has never been elucidated. To address this question, we assembled the largest molecular data set to date, covering > 300 species (out of ca. 800 extant), from 13 of the 14 known families over nine genes, and employed relaxed molecular clocks to uncover their patterns of diversification. As the fossil record of flatfish is contentious, we used sister species distributed on both sides of the American continent to calibrate clock models based on the closure of the Central American Seaway (CAS), and on their current species range. We show that flatfish diversified in two bouts, as species that are today distributed around the Equator diverged during the closure of CAS, while those with a northern range diverged after this, hereby suggesting the existence of a post-CAS closure dispersal for these northern species, most likely along a trans-Arctic northern route, a hypothesis fully compatible with paleogeographic reconstructions.

evolutionary biology

Identifying the genetic determinants of particular phenotypes in microbial genomes with very small training sets

Machine learning (ML) encompasses numerous algorithms that aim at discovering complex patterns between elements within large data using limited prior assumptions or modeling. However, some scientific disciplines still produce small data sets: in particular, empirical studies that try to find the mutations responsible for complex phenotypes are often limited to very small sample sizes (n), while scanning a large number of amino acid sites (p) in a proteome. To date, little is known on how ML performs in this type of so-called \"large p, small n\" problem. To address this question, we evaluated the performance of two general ML classifiers, adaptive boosting (AB) and random forest, on two data sets. To assess the impact of proteome size, we contrasted a small (viral) genome with a larger (bacterial) one. To analyze large proteomes, we further developed a chunking algorithm, and introduce a repeated random forest (RRF) algorithm that stabilizes model predictions. With the influenza data, we were able to rediscover amino acid sites experimentally implicated in three different complex phenotypes (infectivity, transmissibility, and pathogenicity). Results for the larger proteome, pertaining to three types of drug resistance (Ciprofloxacin, Ceftazidime, and Gentamicin), were more nuanced, with RRF making more sensible pre-dictions, with smaller errors rates, than AB. Furthermore, we show that chunking improved runtimes by an order of magnitude and may increase sensitivity of the predictions. Altogether, we demonstrate that ML algorithms can be used to identify genetic determinants in small proteomes (viruses), even with small numbers of individuals. We further show that even if the size of bacterial proteomes pushes AB to its limits in the context of small n, RRF may deserve more scrutiny, which should be facilitated by the plummeting costs of sequencing and, more critically, by phenotyping large cohorts of individuals.\n\nAuthor SummaryFinding the genetic determinants of a phenotype is typically performed by testing for an association between a particular allele and a trait, carrying out the testing over a large number of loci in a large cohort of individuals, itself divided into two subsets of individuals: those who have the trait (cases), and those who do not (controls). However, recruiting large cohorts can be problematic in some experimental fields, while using genotypic information rather than complete genomes can miss some mutations. To address these issues, we implemented two machine learning (ML) algorithms, tweaked for analyzing large genomes and providing stable results. The analysis of a small viral genome, for which genetic determinants of three phenotypes are already known, showed that our approach can rediscover known mutations, almost irrespective of the ML algorithm used. However, the analysis of a larger bacterial genome, for which genetic determinants of three phenotypes are unknown, suggested that the simpler of our modified algorithms performed better, returning more sensitive predictions with lower error rates. This work demonstrates the feasibility of finding genetic determinants of complex phenotypes based on a small number of complete genomes.

bioinformatics

Approximate Bayesian Computation Algorithms for Estimating Network Model Parameters

Studies on Approximate Bayesian Computation (ABC) replacing the intractable likelihood function in evaluation of the posterior distribution have been developed for several years. However, their field of application has to date essentially been limited to inference in population genetics. Here, we propose to extend this approach to estimating the structure of transmission networks of viruses in human populations. In particular, we are interested in estimating the transmission parameters under four very general network structures: random, Watts-Strogatz, Barabasi-Albert and an extension that incorporates aging. Estimation was evaluated under three approaches, based on ABC, ABC-Markov chain Monte Carlo (ABC-MCMC) and ABC-Sequential Monte Carlo (ABC-SMC) samplers. We show that ABC-SMC samplers outperform both ABC and ABC-MCMC, achieving high accuracy and low variance in simulations. This approach paves the way to estimating parameters of real transmission networks of transmissible diseases.

evolutionary biology

Estimation of sub-epidemic dynamics by means of Sequential Monte Carlo Approximate Bayesian Computation: an application to the Swiss HIV Cohort Study

Our ability to accurately infer transmission patterns of infectious diseases is critical to monitor both their spread and the efficacy of public health policies. The use of phylogenetic methods for the reconstruction of viral ancestral relationships has garnered increasing interest, particularly in the characterization of HIV epidemics and sub-epidemics. In the case of this virus, the Swiss HIV Cohort Study (SHCS) contains a wide breadth of genomic data that have been widely used as a means of applying such methods. However, current approaches for quantifying the epidemiological dynamics of diseases are computationally intensive, and fail to scale well with this magnitude of data. To address this issue, we re-implement an Approximate Bayesian Computation (ABC) approach based on sequential Monte Carlo (SMC). By means of simulations, we demonstrate that our implementation is capable of inferring key epidemiological parameters of the Swiss HIV epidemic accurately, and that sampling intensity has no significant effect on the accuracy of our estimates. Applied to a subset of HIV sequences from the SHCS, we show that we can distinguish sub-epidemics that are circulating in culturally distinct Swiss regions. Given these findings, we propose that ABC-SMC samplers will allow us to evaluate the impact of new public health policies, such as the implementation of a needle exchange program in the case of HIV, based on genetic data sampled before and after the implementation of a new policy.

evolutionary biology