bioRxiv Science⌕ Search

Biology subjects

Hauffe, T.

Publications and source records attributed to Hauffe, T..

2 recordsLinked to original sources

Challenges in estimating species age from phylogenetic trees

AimSpecies age, the elapsed time since origination, can give an insight into how species longevity might influence eco-evolutionary dynamics and has been hypothesized to influence extinction risk. Traditionally, species ages have been measured in the fossil record. However, recently, numerous studies have attempted to estimate the ages of extant species from the branch lengths of time-calibrated phylogenies. This approach poses problems because phylogenetic trees contain direct information about species identity only at the tips and not along the branches. Here, we show that incomplete taxon sampling, extinction, and different assumptions about speciation modes can significantly alter the relationship between true species age and phylogenetic branch lengths, leading to high error rates. We found that these biases can lead to erroneous interpretations of eco-evolutionary patterns derived from the comparison between phylogenetic age and other traits, such as extinction risk. InnovationFor bifurcating speciation, which is the default assumption in most analyses, we propose a probabilistic approach to improve the estimation of species ages, based on the properties of a birth-death process. We show that our model can reduce the error by one order of magnitude under cases of high extinction and high percentage of unsampled extant species. Main conclusionOur results call for caution in interpreting the relationship between phylogenetic ages and eco-evolutionary traits, as this can lead to biased and erroneous conclusions. We show that, under the assumption of bifurcate, it is possible to obtain better approximations of species age by combining information from branch lengths with the expectations of a birth-death process.

evolutionary biology↗

Benchmarking imputation methods for discrete biological data

Trait datasets are at the basis of a large share of ecology and evolutionary research, being used to infer ancestral morphologies, to quantify species extinction risks, or to evaluate the functional diversity of biological communities. These datasets, however, are often plagued by missing data, for instance due to incomplete sampling limited data and resource availabilities. Several imputation methods exist to predict missing values and have been successfully evaluated and used to fill the gaps in datasets of quantitative traits. Here we explore the performance of different imputation methods on discrete biological traits i.e. qualitative or categorical traits such as diet or habitat. We develop a bioinformatics pipeline to impute trait data combining phylogenetic, machine learning, and deep learning methods while integrating a simulation framework to evaluate their performance on synthetic datasets. Using this pipeline we run a wide range of simulations under different missing rates, mechanisms, and biases and different evolutionary models. Our results indicate that a new ensemble approach, where we combined the imputation results of a selection of imputation methods provides the most robust and accurate prediction of missing discrete traits. We apply our pipeline to an incomplete trait dataset of 1015 elasmobranch species (including sharks and rays) and found a high imputation accuracy of the predictions based on an expert-based assessment of the missing traits. Our bioinformatic pipeline, implemented in an open-source R package, facilitates the application and comparison of multiple imputation methods to make robust predictions of missing trait values in biological datasets.

bioinformatics↗