bioRxiv Science⌕ Search

Biology subjects

Blanquart, S.

Publications and source records attributed to Blanquart, S..

3 recordsLinked to original sources

AuCoMe: inferring and comparing metabolisms across heterogeneous sets of annotated genomes

Comparative analysis of Genome-Scale Metabolic Networks (GSMNs) may yield important information on the biology, evolution, and adaptation of species. However, it is impeded by the high heterogeneity of the quality and completeness of structural and functional genome annotations, which may bias the results of such comparisons. To address this issue, we developed AuCoMe - a pipeline to automatically reconstruct homogeneous GSMNs from a heterogeneous set of annotated genomes without discarding available manual annotations. We tested AuCoMe with three datasets, one bacterial, one fungal, and one algal, and demonstrated that it successfully reduces technical biases while capturing the metabolic specificities of each organism. Our results also point out shared metabolic traits and divergence points among evolutionarily distant species, such as algae, underlining the potential of AuCoMe to accelerate the broad exploration of metabolic evolution across the tree of life.

systems biology↗

EsMeCaTa: Estimating metabolic capabilities from taxonomic affiliations

PurposeMetabarcoding, and metagenomic sequencing have enabled the characterization of highly diverse environmental communities. The challenge of estimating the metabolic functions carried out by these communities has led to the development of several state-of-the-art methods, most of which are tailored to a specific gene marker. However, the increasing diversity of approaches resulting from advances in sequencing technologies drives the need for methods capable of handling heterogeneous microbial community data. Moreover, predictions often depend on their internal analysis pipelines and are influenced by the underlying databases, which link marker genes to specific functional annotations. This limits users ability to evaluate the quality of predictions by tracing internal data and processes. Finally, users are constrained by the specific annotations provided by these methods (e.g. EC numbers), limiting their ability to conduct further specialized analyses based on intermediate results. MethodsEsMeCaTa predicts consensus proteomes and their associated functions from taxonomic affiliations. A key feature of EsMeCaTa is its explainability and flexibility. To support the flexible integration of heterogeneous sequencing data, EsMeCaTa utilizes taxonomic affiliations obtained through analyses of diverse sequencing datasets. To provide insight into the knowledge available for each taxonomic affliation and to interpret the relevance of predicted functions, EsMeCaTa identifies a taxonomic rank within a given affliation that is suffciently represented by documented proteomes in the UniProt database. The proteins of the UniProt proteomes are clustered and filtered according to a threshold to create consensus proteomes. These consensus proteomes are automatically annotated with functional information (e.g., EC numbers, GO terms) but they are also designed to be used in further customized annotation workflows. Functional annotations are reported in a functional table, which can be enriched with taxon abundances to generate comprehensive functional profiles. ResultsEsMeCaTa predictions have been validated using multiple datasets and compared to a state-of-the-art method. Additionally, it was applied to a novel metabarcoding dataset from a methanogenic reactor, characterizing the microbial community and biogas production across different time points and intake condition. Our results demonstrate the link between biogas production, intake condition and the dynamics of the metabolic functions predicted by EsMeCaTa in the microbial communities.

bioinformatics↗

MATAM: Reconstruction Of Phylogenetic Marker Genes From Short Sequencing Reads In Metagenomes

MotivationAdvances in the sequencing of uncultured environmental samples, dubbed metagenomics, raise a growing need for accurate taxonomic assignment. Accurate identification of organisms present within a community is essential to understanding even the most elementary ecosystems. However, current high-throughput sequencing technologies generate short reads which partially cover full-length marker genes and this poses difficult bioinformatic challenges for taxonomy identification at high resolution\n\nResultsWe designed MATAM, a software dedicated to the fast and accurate targeted assembly of short reads sequenced from a genomic marker of interest. The method implements a stepwise process based on construction and analysis of a read overlap graph. It is applied to the assembly of 16S rRNA markers and is validated on simulated, synthetic and genuine metagenomes. We show that MATAM outperforms other available methods in terms of low error rates and recovered genome fractions and is suitable to provide improved assemblies for precise taxonomic assignments.\n\nAvailabilityhttps://github.com/bonsai-team/matam\n\nContactpierre.pericard@gmail.com, helene.touzet@univ-lille1.fr

bioinformatics↗