bioRxiv Science⌕ Search

Biology subjects

Shi, Z. J.

Publications and source records attributed to Shi, Z. J..

6 recordsLinked to original sources

Maast: genotyping thousands of microbial strains efficiently

Genotyping single nucleotide polymorphisms (SNPs) of intraspecific genomes is a prerequisite to performing population genetic analysis and microbial epidemiology. However, existing algorithms fail to scale for species with thousands of sequenced strains, nor do they account for the biased sampling of strains that has produced considerable redundancy in genome databases. Here we present Maast, a tool that reduces the computational burden of SNP genotyping by leveraging this genomic redundancy. Maast implements a novel algorithm to dynamically identify a minimum set of phylogenetically diverse conspecific genomes that contains the maximum number of SNPs above a user-specified allele frequency. Then it uses these genomes to construct a SNP panel for each species. A species SNP panel enables Maast to rapidly genotype thousands of strains using a hybrid of whole-genome alignment and k-mer exact matching. Maast works with both genome assemblies and unassembled sequencing reads. Compared to existing genotyping methods, Maast is more accurate and up to two orders of magnitude faster. We demonstrate Maasts utility on species with thousands of genomes by reconstructing the genetic structure of Helicobacter pylori across the globe and tracking SARS-CoV-2 diversification during the COVID-19 outbreak. Maast is a fast, reliable SNP genotyping tool that empowers population genetic meta-analysis of microbes at an unrivaled scale. Availabilitysource code of Maast is available at https://github.com/zjshi/Maast. Contactkpollard@gladstone.ucsf.edu

bioinformatics↗

Pitfalls of genotyping microbial communities with rapidly growing genome collections

Detecting genetic variants in metagenomic data is a priority for understanding the evolution, ecology, and functional characteristics of microbial communities. Many recent tools that perform this metagenotyping rely on aligning reads of unknown origin to a reference database of sequences from many species before calling variants. Using simulations designed to represent a wide range of scenarios, we demonstrate that diverse and closely related species both reduce the power and accuracy of reference-based metagenotyping. We identify multi-mapping reads as a prevalent source of errors and illustrate a tradeoff between retaining correct alignments versus limiting incorrect alignments, many of which map reads to the wrong species. Then we quantitatively evaluate several actionable mitigation strategies and review emerging methods with promise to further improve metagenotyping. These findings document a critical challenge that has come to light through the rapid growth of genome collections that push the limits of current alignment algorithms. Our results have implications beyond metagenotyping to the many tools in microbial genomics that depend upon accurate read mapping. HIGHLIGHTSO_LIMost microbial species are genetically diverse. Their single nucleotide variants can be genotyped using metagenomic data aligned to databases constructed from genome collections ("metagenotyping"). C_LIO_LIMicrobial genome collections have grown and now contain many pairs of closely related species. C_LIO_LIClosely related species produce high-scoring but incorrect alignments while also reducing the uniqueness of correct alignments. Both cause metagenotype errors. C_LIO_LIThis dilemma can be mitigated by leveraging paired-end reads, customizing databases to species detected in the sample, and adjusting post-alignment filters. C_LI

bioinformatics↗

EcoFun-MAP: An Ecological Function Oriented Metagenomic Analysis Pipeline

Annotating ecological functions of environmental metagenomes is challenging due to a lack of specialized reference databases and computational barriers. Here we present the Ecological Function oriented Metagenomic Analysis Pipeline (EcoFun-MAP) for efficient analysis of shotgun metagenomes in the context of ecological functions. We manually curated a reference database of EcoFun-MAP which is used for GeoChip design. This database included [~]1,500 functional gene families that were catalogued by important ecological functions, such as carbon, nitrogen, phosphorus, and sulfur cycling, metal homeostasis, stress responses, organic contaminant degradation, antibiotic resistance, microbial defense, electron transfer, virulence and plant growth promotion. EcoFun-MAP has five optional workflows from ultra-fast to ultra-conservative, fitting different research needs from functional gene exploration to stringent comparison. The pipeline is deployed on High Performance Computing (HPC) infrastructure with a highly accessible web-based interface. We showed that EcoFun-MAP is accurate and can process multi-million short reads in a minute. We applied EcoFun-MAP to analyze metagenomes from groundwater samples and revealed interesting insights of microbial functional traits in response to contaminations. EcoFun-MAP is available as a public web server at http://iegst1.rccc.ou.edu:8080/ecofunmap/.

bioinformatics↗

Scalable microbial strain inference in metagenomic data using StrainFacts

While genome databases are nearing a complete catalog of species commonly inhabiting the human gut, their representation of intraspecific diversity is lacking for all but the most abundant and frequently studied taxa. Statistical deconvolution of allele frequencies from shotgun metagenomic data into strain genotypes and relative abundances is a promising approach, but existing methods are limited by computational scalability. Here we introduce StrainFacts, a method for strain deconvolution that enables inference across tens of thousands of metagenomes. We harness a "fuzzy" genotype approximation that makes the underlying graphical model fully differentiable, unlike existing methods. This allows parameter estimates to be optimized with gradient-based methods, speeding up model fitting by two orders of magnitude. A GPU implementation provides additional scalability. Extensive simulations show that StrainFacts can perform strain inference on thousands of metagenomes and has comparable accuracy to more computationally intensive tools. We further validate our strain inferences using single-cell genomic sequencing from a human stool sample. Applying StrainFacts to a collection of more than 10,000 publicly available human stool metagenomes, we quantify patterns of strain diversity, biogeography, and linkage-disequilibrium that agree with and expand on what is known based on existing reference genomes. StrainFacts paves the way for large-scale biogeography and population genetic studies of microbiomes using metagenomic data.

bioinformatics↗

Ultra-rapid metagenotyping of the human gut microbiome

Sequence variation is used to quantify population structure and identify genetic determinants of phenotypes that vary within species. In the human microbiome and other environments, single nucleotide polymorphisms (SNPs) are frequently detected by aligning metagenomic sequencing reads to catalogs of genes or genomes. But this requires high-performance computing and enough read coverage to distinguish SNPs from sequencing errors. We solved these problems by developing the GenoTyper for Prokaytotes (GT-Pro), a suite of novel methods to catalog SNPs from genomes and use exact k-mer matches to perform ultra-fast reference-based SNP calling from metagenomes. Compared to read alignment, GT-Pro is more accurate and two orders of magnitude faster. We discovered 104 million SNPs in 909 human gut species, characterized their global population structure, and tracked pathogenic strains. GT-Pro democratizes strain-level microbiome analysis by making it possible to genotype hundreds of metagenomes on a personal computer. Software availabilityGT-Pro is available at https://github.com/zjshi/gt-pro.

bioinformatics↗

A unified sequence catalogue of over 280,000 genomes obtained from the human gut microbiome

Comprehensive reference data is essential for accurate taxonomic and functional characterization of the human gut microbiome. Here we present the Unified Human Gastrointestinal Genome (UHGG) collection, a resource combining 286,997 genomes representing 4,644 prokaryotic species from the human gut. These genomes contain over 625 million protein sequences used to generate the Unified Human Gastrointestinal Protein (UHGP) catalogue, a collection that more than doubles the number of gut protein clusters over the Integrated Gene Catalogue. We find that a large portion of the human gut microbiome remains to be fully explored, with over 70% of the UHGG species lacking cultured representatives, and 40% of the UHGP missing meaningful functional annotations. Intra-species genomic variation analyses revealed a large reservoir of accessory genes and single-nucleotide variants, many of which were specific to individual human populations. These freely available genomic resources should greatly facilitate investigations into the human gut microbiome.

microbiology↗