bioRxiv ScienceSearch

Biology subjects

Roach, J.

Publications and source records attributed to Roach, J..

2 recordsLinked to original sources

Distribution-based comprehensive evaluation ofmethods for differential expression analysis inmetatranscriptomics

Understanding the function of the human microbiome is important; however, the development of statistical methods specifically for the microbial gene expression (i.e., metatranscriptomics) is in its infancy. Many currently employed differential expression analysis methods have been designed for different data types and have not been evaluated in metatranscriptomics settings. To address this gap, we undertook a comprehensive evaluation and benchmarking of ten differential analysis methods for metatranscriptomics data. We used a combination of real and simulated data to evaluate performance (i.e., model fit, type I error, false discovery rate, and sensitivity) of the methods: log-normal (LN), logistic-beta (LB), MAST, DESeq2, metagenomeSeq, ANCOM-BC, LEfSe, ALDEx2, Kruskal-Wallis, and two-part Kruskal-Wallis. The simulation was informed by supragingival biofilm microbiome data from 300 preschool-age children enrolled in a study of early childhood caries (ECC), whereas validations were sought in two additional datasets from an ECC study and an inflammatory bowel disease (IBD) study. The LB test showed the highest sensitivity in both small and large samples and reasonably controlled type I error. Contrarily, MAST was hampered by inflated type I error. Upon application of the LN and LB tests in the ECC study, we found that genes C8PHV7 and C8PEV7, harbored by the lactate-producing Campylobacter gracilis, had the strongest association with childhood dental diseases. This comprehensive model evaluation offer practical guidance for selection of appropriate methods for rigorous analyses of differential expression in metatranscriptomics. Selection of an optimal method increases the possibility of detecting true signals while minimizing the chance of claiming false ones.

bioinformatics

Improved Metabolite Prediction Using Microbiome Data-Based Elastic Net Models

Microbiome data are becoming increasingly available in large health cohorts yet metabolomics data are still scant. While many studies generate microbiome data, they lack matched metabolomics data or have considerable missing proportions of metabolites. Since metabolomics is key to understanding microbial and general biological activities, the possibility of imputing individual metabolites or inferring metabolomics pathways from microbial taxonomy or metagenomics is intriguing. Importantly, current metabolomics profiling methods such as the HMP Unified Metabolic Analysis Network (HUMAnN) have unknown accuracy and are limited in their ability to predict individual metabolites. To address this gap, we developed a novel metabolite prediction method, and we present its application and evaluation in an oral microbiome study. We developed ENVIM based on the Elastic Net Model (ENM) to predict metabolites using micorbiome data. ENVIM introduces an extra step to ENM to consider variable importance scores and thus achieve better prediction power. We investigate the metabolite prediction performance of ENVIM using metagenomic and metatranscriptomic data in a supragingival biofilm multi-omics dataset of 297 children ages 3-5 who were participants of a community-based study of early childhood oral health (ZOE 2.0) in North Carolina, United States. We further validate ENVIM in two additional publicly available multi-omics datasets generated from studies of gut health and vagina health. We select gene-family sets based on variable importance scores and modify the existing ENM strategy used in the MelonnPan prediction software to accommodate the unique features of microbiome and metabolome data. We evaluate metagenomic and metatranscriptomic predictors and compare the prediction performance of ENVIM to the standard ENM employed in MelonnPan. The newly-developed ENVIM method showed superior metabolite predictive accuracy than MelonnPan using metatranscriptomics data only, metagenomics data only, or both of these two. Both methods perform better prediction using gut or vagina microbiome data than using oral microbiome data for the samples corresponding metabolites. The top predictable compounds have been reported in all these three datasets from three different body sites. Enrichment of prediction some contributing species has been detected.

microbiology