bioRxiv Science⌕ Search

Biology subjects

Kav, A. B.

Publications and source records attributed to Kav, A. B..

2 recordsLinked to original sources

Comparative metagenomics using pan-metagenomics graphs

Identifying microbial genomic factors underlying human phenotypes is a key goal of microbiome research. Sequence graphs are a highly effective tool for genome comparisons because they enable high-resolution de novo analyses that capture and contextualize complex genomic variation. However, applying sequence graphs to complex microbial communities remains challenging due to the scale and complexity of metagenomic data. Existing multi-sample sequence graphs used in these settings are highly complex, computationally expensive, less accurate than single-sample alternatives, and often involve arbitrary coarse-graining. Here, we present copangraph, a multi-sample sequence-graph-based analysis framework for comprehensive comparisons of genomic variation across metagenomes. Copangraph uses a novel homology-based graph, which provides both non-arbitrary, evolutionary-motivated grouping of sequences into the same node as well as flexibility in the scale of variation represented by the graph. Its construction relies on hybrid coassembly, a new coassembly approach in which single-sample graphs are first constructed separately and are then merged to create a multi-sample graph. We also present an algorithm that uses paired-end reads to improve detection of contiguous genomic regions, increasing accuracy. Our results demonstrate that copangraph captures sequence and variant information more accurately than alternative methods, provides graphs that are more suitable for comparative analysis than de Bruijn graphs, and is computationally tractable. We show that copangraph reflects meaningful metagenomic variation across diverse scenarios. Importantly, it enables significantly better performance than other metagenomic representations when predicting the gut colonization trajectories of Vancomycin-resistant Enterococcus. Our results underscore the value of our multi-sample, graph-based framework for comparative metagenomic analyses.

bioinformatics↗

Identification of Sample Processing Errors in Microbiome Studies Using Host Genetic Profiles

In microbiome studies, sample processing errors are frequent and difficult to detect, especially in large studies involving multiple sites, personnel, and sample types. We present two complementary approaches to identify such errors using host DNA profiled via metagenomic sequencing of microbiome samples. The first approach compares host SNPs inferred from metagenomics to independently obtained genotypes (e.g., microarray genotypes) to match samples to their donors, while the second method compares metagenomics-inferred SNPs between samples to identify samples supplied by the same donor. Furthermore, we demonstrate that combining these methods with experimental metadata provides greater confidence in the identification of errors. Analyzing a longitudinal vaginal microbiome dataset, we demonstrate the ability of our approach to identify mislabeled samples. Using subsampling, we further show that our methods are robust to low sequencing coverage. Overall, our analysis highlights the frequency of processing errors in microbiome studies. We therefore recommend applying error-detection methods in all studies with suitable data.

bioinformatics↗