bioRxiv Science⌕ Search

Biology subjects

Correa, F. B.

Publications and source records attributed to Correa, F. B..

2 recordsLinked to original sources

MuDoGeR: Multi-Domain Genome Recovery from metagenomes made easy

Several computational frameworks and workflows that recover genomes from prokaryotes, eukaryotes, and viruses from metagenomes exist. However, it is difficult for scientists with little bioinformatics experience to evaluate quality, annotate genes, dereplicate, assign taxonomy and calculate relative abundance and coverage of genomes belonging to different domains. MuDoGeR is a user-friendly tool accessible for non-bioinformaticians that make it easy to recover genomes of prokaryotes, eukaryotes, and viruses from metagenomes, either alone or in combination. We tested MuDoGer using 24 individual-isolated genomes and 574 metagenomes, demonstrating the applicability for a few samples and high throughput. MuDoGeR is open-source software available at https://github.com/mdsufz/MuDoGeR.

bioinformatics↗

TAG.ME: Taxonomic Assignment of Genetic Markers for Ecology

1.BackgroundSequencing of amplified genetic markers, such as the 16S rRNA gene, have been extensively used to characterize microbial community composition. Recent studies suggested that Amplicon Sequences Variants (ASV) should replace the Operational Taxonomic Units (OTU), given the arbitrary definition of sequence identity thresholds used to define units. Alignment-free methods are an interesting alternative for the taxonomic classification of the ASVs, preventing the introduction of biases from sequence identity thresholds.\n\nResultsHere we present TAG.ME, a novel alignment-independent and amplicon-specific method for taxonomic assignment based on genetic markers. TAG.ME uses a multilevel supervised learning approach to create predictive models based on user-defined genetic marker genes. The predictive method can assign taxonomy to sequenced amplicons efficiently and effectively. We applied our method to assess gut and soil sample classification, and it outperformed alternative approaches, identifying a substantially larger proportion of species. Benchmark tests performed using the RDP database, and Mock communities reinforced the precise classification into deep taxonomic levels.\n\nConclusionTAG.ME presents a new approach to assign taxonomy to amplicon sequences accurately. Our classification model, trained with amplicon specific sequences, can address resolution issues not solved by other methods and approaches that use the whole 16S rRNA gene sequence. TAG.ME is implemented as an R package and is freely available at http://gabrielrfernandes.github.io/tagme/

bioinformatics↗