bioRxiv Science⌕ Search

Biology subjects

Karwowska, Z.

Publications and source records attributed to Karwowska, Z..

3 recordsLinked to original sources

Effects of Data Transformation and Model Selection on Feature Importance in Microbiome Classification Data

Accurate classification of host phenotypes from microbiome data is essential for future therapies in microbiome-based medicine and machine learning approaches have proved to be an effective solution for the task. The complex nature of the gut microbiome, data sparsity, compositionality and population-specificity however remain challenging, which highlights the critical need for standardized methodologies to improve the accuracy and reproducibility of the results. Microbiome data transformations can alleviate some of the aforementioned challenges, but their usage in machine learning tasks has largely been unexplored. Our aim was to assess the impact of various data transformations on the accuracy, generalizability and feature selection by analysis using more than 8,500 samples from 24 shotgun metagenomic datasets. Our findings demonstrate the feasibility of distinguishing between healthy and diseased individuals using microbiome data with minimal dependence on the algorithm and transformation selection. Remarkably, presence-absence transformation performed comparably well to abundance-based transformations, and only a small subset of predictors is crucial for accurate classification. However, while different transformations resulted in comparable classification performance, the most important features varied significantly, which highlight the need to reevaluate machine-learning based biomarker detection. Our research provides valuable guidance for applying machine learning on microbiome data, offering novel insights and highlighting important areas for future research.

bioinformatics↗

Microbiome time series data reveal predictable patterns of change

BackgroundThe gut microbiome is crucial for human health and disease. Longitudinal studies are gaining importance in understanding its dynamics over time, compared to cross-sectional approaches. Investigating the temporal dynamics of the microbiome, including individual bacterial species and clusters, is essential for comprehending its functionality and impact on health. This knowledge has implications for targeted therapeutic strategies, such as personalized diets and probiotic therapy. ResultsHere, by adopting a rigorous statistical approach, we aim to shed light on the temporal changes in the gut microbiome and unravel its intricate behavior over time. We leveraged four long and dense time series of the gut microbiome in generally healthy individuals examining how its composition evolves as a community and how individual bacterial species behave over time. We also explore whether specific clusters of bacteria exhibit similar fluctuations, which could provide insights into potential functional relationships and interactions within the microbiome Our study reveals that despite its high volatility, the human gut microbiome is stable in time and can be predicted based solely on its previous states. We characterize the unique temporal behavior of individual bacterial species and identify distinct longitudinal regimes in which bacteria exhibit specific patterns of behavior. Finally, through cluster analysis, we identify groups of bacteria that exhibit coordinated fluctuations over time. ConclusionsOur findings contribute to our understanding of the dynamic nature of the gut microbiome and its potential implications for human health. The provided guidelines support scientists studying gut microbiome complex dynamics, promoting further research and advancements in microbiome analysis.

bioinformatics↗

Deep embeddings to comprehend and visualize microbiome protein space

Understanding the function of microbial proteins is essential to reveal the clinical potential of the microbiome. The application of high-throughput sequencing technologies allows for fast and increasingly cheaper acquisition of data from microbial communities. However, many of the inferred protein sequences are novel and not catalogued, hence the possibility of predicting their function through conventional homology-based approaches is limited. Here, we leverage a deep-learning-based representation of proteins to assess its utility in alignment-free analysis of microbial proteins. We trained a language model on the Unified Human Gastrointestinal Protein catalogue and validated the resulting protein representation on the bacterial part of the SwissProt database. Finally, we present a use case on proteins involved in SCFA metabolism. Results indicate that the deep learning model manages to accurately represent features related to protein structure and function, allowing for alignment-free protein analyses. Technologies that contextualize metagenomic data are a promising direction to deeply understand the microbiome.

bioinformatics↗