bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.07.21.549984

Looking at the Full Picture: Utilizing Topic Modeling to Determine Disease-Associated Microbiome Communities

Abstract

The microbiome is a complex micro-ecosystem that provides the host with pathogen defense, food metabolism, and other vital processes. Alterations of the microbiome (dysbiosis) have been linked with a number of diseases such as cancers, multiple sclerosis (MS), Alzheimers disease, etc. Generally, differential abundance testing between the healthy and patient groups is performed to identify important bacteria (enriched or depleted in one group). However, simply providing a singular species of bacteria to an individual lacking that species for health improvement has not been as successful as fecal matter transplant (FMT) therapy. Interestingly, FMT therapy transfers the entire gut microbiome of a healthy (or mixture of) individual to an individual with a disease. FMTs do, however, have limited success, possibly due to concerns that not all bacteria in the community may be responsible for the healthy phenotype. Therefore, it is important to identify the community of microorganisms linked to the health as well as the disease state of the host. Here we applied topic modeling, a natural language processing tool, to assess latent interactions occurring among microbes; thus, providing a representation of the community of bacteria relevant to healthy vs. disease state. Specifically, we utilized our previously published data that studied the gut microbiome of patients with relapsing-remitting MS (RRMS), a neurodegenerative autoimmune disease that has been linked to a variety of factors, including a dysbiotic gut microbiome. With topic modeling we identified communities of bacteria associated with RRMS, including genera previously discovered, but also other taxa that would have been overlooked simply with differential abundance testing. Our work shows that topic modeling can be a useful tool for analyzing the microbiome in dysbiosis and that it could be considered along with the commonly utilized differential abundance tests to better understand the role of the gut microbiome in health and disease. Author SummaryTrillion of bacteria (microbiome) living in and on the human body play an important role in keeping us healthy and an alteration in their composition has been linked to multiple diseases such as cancers, multiple sclerosis (MS), and Alzheimers. Identifying specific bacteria for targeted therapies is crucial, however studying individual bacteria fails to capture their interactions within the microbial community. The relative success of fecal matter transplants (FMTs) from healthy individual(s) to patients and the failure of individual bacterial therapy suggests the importance of the microbiome community in health. Therefore, there is a need to develop tools to identify the communities of microbes making up the healthy and disease state microbiome. Here we applied topic modeling, a natural language processing tool, to identify microbial communities associated with relapsing-remitting MS (RRMS). Specifically, we show the advantage of topic modeling in identifying the bacterial community structure of RRMS patients, which includes previously reported bacteria linked to RRMS but also otherwise overlooked bacteria. These results reveal that integrating topic modeling with traditional approaches improves the understanding of the microbiome in RRMS and it could be employed with other diseases that are known to have an altered microbiome.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shrode, R. L., Ollberding, N. J., Mangalam, A. K.. 2023-07-25. Looking at the Full Picture: Utilizing Topic Modeling to Determine Disease-Associated Microbiome Communities. https://doi.org/10.1101/2023.07.21.549984

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

spatialMET: an open and scalable framework for spatial metabolomics analysis

Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.

bioinformatics↗

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics↗

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics↗