bioRxiv Science⌕ Search

Biology subjects

Sepich-Poore, G. D.

Publications and source records attributed to Sepich-Poore, G. D..

3 recordsLinked to original sources

Reply to: Caution Regarding the Specificities of Pan-Cancer Microbial Structure

The cancer microbiome field tremendously accelerated following the release of our manuscript nearly three years ago1, including direct validation of our cancer type-specific conclusions in independent, international cohorts2,3 and the tumor microbiomes adoption into the hallmarks of cancer4. Disentangling contamination signals from biological signals is an important consideration for this research field. Therefore, despite numerous, high-impact, peer-reviewed research papers that either validated our conclusions or extended them using data we released2,5-13, we carefully considered criticism raised by Gihawi et al. about potential mishandling of contaminants, batch effects, and machine learning approaches--all of which were central topics in our manuscript. Nonetheless, a close examination of each concern alongside the original manuscript and re-analyses of our published data strongly demonstrates the robustness of the original findings. To remove all doubt, however, we have reproduced all key conclusions from the original manuscript using only overlapping bacterial genera identified in a highly decontaminated, multi-cancer, international cohort (Weizmann Institute of Science, WIS)2, with or without batch correction, and with multiclass machine learning analyses to mitigate class imbalances. Our published pan-cancer mycobiome manuscript3 also affirms these findings using updated, state-of-the-art methods. We also note that every analysis shown here was possible using public data and code that we had already provided.

bioinformatics↗

BIRDMAn: A Bayesian differential abundance framework that enables robust inference of host-microbe associations

Quantifying the differential abundance (DA) of specific taxa among experimental groups in microbiome studies is challenging due to data characteristics (e.g., compositionality, sparsity) and specific study designs (e.g., repeated measures, meta-analysis, cross-over). Here we present BIRDMAn (Bayesian Inferential Regression for Differential Microbiome Analysis), a flexible DA method that can account for microbiome data characteristics and diverse experimental designs. Simulations show that BIRDMAn models are robust to uneven sequencing depth and provide a >20-fold improvement in statistical power over existing methods. We then use BIRDMAn to identify antibiotic-mediated perturbations undetected by other DA methods due to subject-level heterogeneity. Finally, we demonstrate how BIRDMAn can construct state-of-the-art cancer-type classifiers using The Cancer Genome Atlas (TCGA) dataset, with substantial accuracy improvements over random forests and existing DA tools across multiple sequencing centers. Collectively, BIRDMAn extracts more informative biological signals while accounting for study-specific experimental conditions than existing approaches.

bioinformatics↗

OGUs enable effective, phylogeny-aware analysis of even shallow metagenome community structures

We introduce Operational Genomic Unit (OGU), a metagenome analysis strategy that directly exploits sequence alignment hits to individual reference genomes as the minimum unit for assessing the diversity of microbial communities and their relevance to environmental factors. This approach is independent from taxonomic classification, granting the possibility of maximal resolution of community composition, and organizes features into an accurate hierarchy using a phylogenomic tree. The outputs are suitable for contemporary analytical protocols for community ecology, differential abundance and supervised learning while supporting phylogenetic methods, such as UniFrac and phylofactorization, that are seldomly applied to shotgun metagenomics despite being prevalent in 16S rRNA gene amplicon studies. As demonstrated in one synthetic and two real-world case studies, the OGU method produces biologically meaningful patterns from microbiome datasets. Such patterns further remain detectable at very low metagenomic sequencing depths. Compared with taxonomic unit-based analyses implemented in currently adopted metagenomics tools, and the analysis of 16S rRNA gene amplicon sequence variants, this method shows superiority in informing biologically relevant insights, including stronger correlation with body environment and host sex on the Human Microbiome Project dataset, and more accurate prediction of human age by the gut microbiomes in the Finnish population. We provide Woltka, a bioinformatics tool to implement this method, with full integration with the QIIME 2 package and the Qiita web platform, to facilitate OGU adoption in future metagenomics studies. ImportanceShotgun metagenomics is a powerful, yet computationally challenging, technique compared to 16S rRNA gene amplicon sequencing for decoding the composition and structure of microbial communities. However, current analyses of metagenomic data are primarily based on taxonomic classification, which is limited in feature resolution compared to 16S rRNA amplicon sequence variant analysis. To solve these challenges, we introduce Operational Genomic Units (OGUs), which are the individual reference genomes derived from sequence alignment results, without further assigning them taxonomy. The OGU method advances current read-based metagenomics in two dimensions: (i) providing maximal resolution of community composition while (ii) permitting use of phylogeny-aware tools. Our analysis of real-world datasets shows several advantages over currently adopted metagenomic analysis methods and the finest-grained 16S rRNA analysis methods in predicting biological traits. We thus propose the adoption of OGU as standard practice in metagenomic studies.

bioinformatics↗