bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.04.12.535864

A comparison between Greengenes, SILVA, RDP, and NCBI reference databases in four published microbiota datasets

Abstract

Inaccurate bacterial taxonomic assignment in 16S-based microbiota experiments could have deleterious effects on research results, as all downstream analyses heavily rely on the accurate assessment of microbial taxonomy: a bias in the choice of the reference database can deeply alter microbiota biodiversity (alpha-diversity), composition (beta-diversity), and taxa profile (bacterial relative abundances). In this paper, we explored the influence of the reference 16S rRNA collection by performing a classification against four of the main databases used by the scientific community (i.e. Greengenes, SILVA, RDP, NCBI); the consequences of database clustering at 97% were also explored. To investigate the effects of the database choice on real and representative microbiome samples from different ecosystems, we performed a comparative analysis on four already published datasets from various sources: stools from a mouse model experiment, bovine milk, human gut microbiota stool samples, and swabs from the human vaginal environment. We took into consideration the computational time needed to perform the taxonomic classification as well. Although values in both alpha- and beta-diversity varied a lot, sometimes even statistically, according to the dataset chosen and the eventual clustering, the final outcome of the analysis was a concordance in the capability to retrieve the original experimental group differences over the various datasets. However, in the taxonomy classification, we found several inconsistencies with taxonomies correctly assigned in only some of the four databases. The degree of concordance among the databases was related to both the complexity of the environment and its degree of completeness in the reference databases. IMPORTANCE16S rRNA sequencing is, nowadays, the most commonly used strategy for microbiota profiling in many different ecosystems, ranging from human-associated to animal models, food matrices, and environmental samples. The ability of this kind of analysis to correctly capture differences in the microbiota composition is related to the taxonomic classification of the fragments obtained from sequencing and, thus, to the choice of the best reference database. This paper deals with four of the most popular microbial databases, which were evaluated in their ability to reproduce the experimental evidence from four already published datasets. The knowledge of the advantages and drawbacks of the database choice can be pivotal for planning future experiments in the field, making researchers aware of the repercussions of such a choice according to the different environments under scrutiny. Moreover, this work can also shed new light upon past results, partially explaining discordant evidence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ceccarani, C., Severgnini, M.. 2023-04-13. A comparison between Greengenes, SILVA, RDP, and NCBI reference databases in four published microbiota datasets. https://doi.org/10.1101/2023.04.12.535864

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

spatialMET: an open and scalable framework for spatial metabolomics analysis

Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.

bioinformatics↗

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics↗

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics↗