bioRxiv Science⌕ Search

Biology subjects

Austin, G. I.

Publications and source records attributed to Austin, G. I..

3 recordsLinked to original sources

Early prediction of preeclampsia using the first trimester vaginal microbiome

Preeclampsia is a severe obstetrical syndrome which contributes to 10-15% of all maternal deaths. Although the mechanisms underlying systemic damage in preeclampsia--such as impaired placentation, endothelial dysfunction, and immune dysregulation--are well studied, the initial triggers of the condition remain largely unknown. Furthermore, although the pathogenesis of preeclampsia begins early in pregnancy, there are no early diagnostics for this life-threatening syndrome, which is typically diagnosed much later, after systemic damage has already manifested. Here, we performed deep metagenomic sequencing and multiplex immunoassays of vaginal samples collected during the first trimester from 124 pregnant individuals, including 62 who developed preeclampsia with severe features. We identified multiple significant associations between vaginal immune factors, microbes, clinical factors, and the early pathogenesis of preeclampsia. These associations vary with BMI, and stratification revealed strong associations between preeclampsia and Bifidobacterium spp., Prevotella timonensis, and Sneathia vaginalis. Finally, we developed machine learning models that predict the development of preeclampsia using this first trimester data, collected ~5.7 months prior to clinical diagnosis, with an auROC of 0.78. We validated our models using data from an independent cohort (MOMS-PI), achieving an auROC of 0.80. Our findings highlight robust associations among the vaginal microbiome, local host immunity, and early pathogenic processes of preeclampsia, paving the way for early detection, prevention and intervention for this devastating condition.

microbiology↗

Compositional transformations can reasonably introduce phenotype-associated values into sparse features

It was recently argued1 that an analysis of tumor-associated microbiome data2 is invalid because features that were originally very sparse (genera with mostly zero read counts) became associated with the phenotype following batch correction1. Here, we examine whether such an observation should necessarily indicate issues with processing or machine learning pipelines. We show counterexamples using the centered log ratio (CLR) transformation, which is often used for analysis of compositional microbiome data3. The CLR transformation has similarities to Voom-SNM4,5, the batch-correction method brought into question1,2, yet is a sample-wise operation that cannot, in itself, "leak" information or invalidate downstream analyses. We show that because the CLR transformation divides each value by the geometric mean of its sample, common imputation strategies for missing or zero values result in transformed features that are associated with the geometric mean. Through analyses of both synthetic and vaginal microbiome datasets we demonstrate that when the geometric mean is associated with a phenotype, sparse and CLR-transformed features will also become associated with it. We re-analyze features highlighted by Gihawi et al.1 and demonstrate that the phenomenon of sparse features becoming phenotype-associated can also be observed after a CLR transformation, which serves as a counterexample to the claim that such an observation necessarily means information leakage. While we do not intend to address other concerns regarding tumor microbiome analyses1,6, validate Poore et al.s2 results, or evaluate batch-correction pipelines, we conclude that because phenotype-associated features that were initially sparse can be created by a sample-wise transformation that cannot artifactually inflate machine learning performance, their detection is not independently sufficient to demonstrate information leakage in machine learning pipelines. Microbiome data is multivariate, and as such, a value of zero carries a different meaning for each sample. Many transformations, including CLR and other batch-correction methods, are likewise multivariate, and, as these issues demonstrate, each individual feature should be interpreted with caution.

bioinformatics↗

Processing-bias correction with DEBIAS-M improves cross-study generalization of microbiome-based prediction models

Every step in common microbiome profiling protocols has variable efficiency for each microbe. For example, different DNA extraction kits may have different efficiency for Gram-positive and -negative bacteria. These variable efficiencies, combined with technical variation, create strong processing biases, which impede the identification of signals that are reproducible across studies and the development of generalizable and biologically interpretable prediction models. "Batch-correction" methods have been used to alleviate these issues computationally with some success. However, many make strong parametric assumptions which do not necessarily apply to microbiome data or processing biases, or require the use of an outcome variable, which risks overfitting. Lastly and importantly, existing transformations used to correct microbiome data are largely non-interpretable, and could, for example, introduce values to features that were initially mostly zeros. Altogether, processing bias currently compromises our ability to glean robust and generalizable biological insights from microbiome data. Here, we present DEBIAS-M (Domain adaptation with phenotype Estimation and Batch Integration Across Studies of the Microbiome), an interpretable framework for inference and correction of processing bias, which facilitates domain adaptation in microbiome studies. DEBIAS-M learns bias-correction factors for each microbe in each batch that simultaneously minimize batch effects and maximize cross-study associations with phenotypes. Using benchmarks of HIV and colorectal cancer classification from gut microbiome data, and cervical neoplasia prediction from cervical microbiome data, we demonstrate that DEBIAS-M outperforms batch-correction methods commonly used in the field. Notably, we show that the inferred bias-correction factors are stable, interpretable, and strongly associated with specific experimental protocols. Overall, we show that DEBIAS-M allows for better modeling of microbiome data and identification of interpretable signals that are reproducible across studies.

bioinformatics↗