bioRxiv Science⌕ Search

Biology subjects

Jaing, C.

Publications and source records attributed to Jaing, C..

4 recordsLinked to original sources

An Embeddings Fusion Approach Predicts Disease State from Microbiome Features

BackgroundDeep neural networks are a proven technique for working with high dimensional data because of their ability to draw-out meaningful patterns and create vector representations known as "embeddings", which make it easier to work with learning tasks on large inputs as they capture the semantics and variance of the data. Microbial community abundance profiles are well suited for an embeddings approach due to their high dimensionality, and in this work we introduce a novel approach for generating embeddings from visual representations which encode NCBIs taxonomic tree and microbial compositions as images, enabling the creation of embeddings that capture factors such as disease status, type, and geographical location. ResultsWe profiled 13,534 public human metagenomes spanning 85 studies, 24 disease types, 35 countries, and 31,756 microbial species using a profiling pipeline that indexes NCBIs nucleotide database (nt) across all kingdoms of life. Our model achieves an average classification performance of 84% in distinguishing healthy and disease conditions; 87% for disease types, 99% for body sites, and 88% for geographical locations. It also achieves a 97% accuracy when performing multi-label classification of the four factors combined. ConclusionOur work highlights the use of an embeddings approach that can encode multiple features and create efficient contextualization of profiled metagenomes derived from microbiome samples. The models embeddings can be used to cluster existing samples based on multiple conditions and interpretations, and new embeddings can be quickly created for new samples and fitted to existing clusters to characterize them. This has practical applications for unknown, unlabeled microbiome samples.

bioinformatics↗

Beyond Microbial Abundance: Metadata Integration Enhances Disease Prediction in Human Microbiome Studies

Multiple studies have highlighted the human microbiomes potential as a biomarker for diagnosing diseases through its interaction with systems like the gut, immune, liver, and skin via key axes. Advances in sequencing technologies and highperformance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the host and sample collection/processing protocols. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a common pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising of 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

microbiology↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13,897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on Jan 04, 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

microbiology↗

Engineering immunogens that select for specific mutations in HIV broadly neutralizing antibodies

Vaccine development targeting rapidly evolving pathogens such as HIV-1 requires induction of broadly neutralizing antibodies (bnAbs) with conserved paratopes and mutations, and, in some cases, the same Ig-heavy chains. The current trial-and-error search for immunogen modifications that improve selection for specific bnAb mutations is imprecise. To precisely engineer bnAb boosting immunogens, we used molecular dynamics simulations to examine encounter states that form when antibodies collide with the HIV-1 Envelope (Env). By mapping how bnAbs use encounter states to find their bound states, we identified Env mutations that were predicted to select for specific antibody mutations in two HIV-1 bnAb B cell lineages. The Env mutations encoded antibody affinity gains and selected for desired antibody mutations in vivo. These results demonstrate proof-of-concept that Env immunogens can be designed to directly select for specific antibody mutations at residue-level precision by vaccination, thus demonstrating the feasibility of sequential bnAb-inducing HIV-1 vaccine design.

immunology↗