bioRxiv ScienceSearch

Biology subjects

Berger, B.

Publications and source records attributed to Berger, B..

6 recordsLinked to original sources

Carnelian: alignment-free functional binning and abundance estimation of metagenomic reads

Accurate assignment of metagenomic reads to their functional roles is an important first step towards gaining insights into the relationship between the human microbiomeincluding the collective genesand disease. Existing approaches focus on binning sequencing reads into known taxonomic classes or by genes, often failing to produce results that generalize across different cohorts with the same disease. We present Carnelian, a highly precise and accurate pipeline for alignment-free functional binning and abundance estimation, which leverages the recent idea of even-coverage, low-density locality sensitive hashing. When coupled with one-against-all classifiers, reads can be binned by molecular function encoded in their gene content with higher precision and accuracy. Carnelians minutes-per-metagenome processing speed enables analysis of large-scale disease or environmental datasets to reveal disease- and environment-specific changes in microbial functionality previously poorly understood. Our pipeline newly reveals a functional dysbiosis in patient gut microbiomes, not found in earlier metagenomic studies, and identifies a distinct shift from matched healthy individuals in Type-2 Diabetes (T2D) and early-stage Parkinsons Disease (PD). We remarkably identify a set of functional markers that can differentiate between patients and healthy individuals consistently across both the datasets with high specificity.

bioinformatics

Panoramic stitching of heterogeneous single-cell transcriptomic data

Researchers are generating single-cell RNA sequencing (scRNA-seq) profiles of diverse biological systems1-4 and every cell type in the human body.5 Leveraging this data to gain unprecedented insight into biology and disease will require assembling heterogeneous cell populations across multiple experiments, laboratories, and technologies. Although methods for scRNA-seq data integration exist6,7, they often naively merge data sets together even when the data sets have no cell types in common, leading to results that do not correspond to real biological patterns. Here we present Scanorama, inspired by algorithms for panorama stitching, that overcomes the limitations of existing methods to enable accurate, heterogeneous scRNA-seq data set integration. Our strategy identifies and merges the shared cell types among all pairs of data sets and is orders of magnitude faster than existing techniques. We use Scanorama to combine 105,476 cells from 26 diverse scRNA-seq experiments across 9 different technologies into a single comprehensive reference, demonstrating how Scanorama can be used to obtain a more complete picture of cellular function across a wide range of scRNA-seq experiments.

bioinformatics

Neural Data Visualization for Scalable and Generalizable Single Cell Analysis

Single-cell RNA sequencing is becoming effective and accessible as emerging technologies push its scale to millions of cells and beyond. Visualizing the landscape of single cell expression has been a fundamental tool in single cell analysis. However, standard methods for visualization, such as t-stochastic neighbor embedding (t-SNE), not only lack scalability to data sets with millions of cells, but also are unable to generalize to new cells, an important ability for transferring knowledge across fast-accumulating data sets. We introduce net-SNE, which trains a neural network to learn a high quality visualization of single cells that newly generalizes to unseen data. While matching the visualization quality of t-SNE on 14 benchmark data sets of varying sizes, from hundreds to 1.3 million cells, net-SNE also effectively positions previously unseen cells, even when an entire subtype is missing from the initial data set or when the new cells are from a different sequencing experiment. Furthermore, given a \"reference\" visualization, net-SNE can vastly reduce the computational burden of visualizing millions of single cells from multiple days to just a few minutes of runtime. Our work provides a general framework for newly bootstrapping single cell analysis from existing data sets.

bioinformatics

Latent variable model for aligning barcoded short-reads improves downstream analyses

Recent years have seen the emergence of several \"third-generation\" sequencing platforms, each of which aims to address shortcomings of standard next-generation short-read sequencing by producing data that capture long-range information, thereby allowing us to access regions of the genome that are inaccessible with short-reads alone. These technologies either produce physically longer reads typically with higher error rates or instead capture long-range information at low error rates by virtue of read \"barcodes\" as in 10x Genomics Chromium platform. As with virtually all sequencing data, sequence alignment for third-generation sequencing data is the foundation on which all downstream analyses are based. Here we introduce a latent variable model for improving barcoded read alignment, thereby enabling improved downstream genotyping and phasing. We demonstrate the feasibility of this approach through developing EMerAld-- or EMA for short-- and testing it on the barcoded short-reads produced by 10xs sequencing technologies. EMA not only produces more accurate alignments, but unlike other methods also assigns interpretable probabilities to the alignments it generates. We show that genotypes called from EMAs alignments contain over 30% fewer false positives than those called from Lariats (the current 10x alignment tool), with a fewer number of false negatives, on datasets of NA12878 and NA24385 as compared to NIST GIAB gold standard variant calls. Moreover, we demonstrate that EMA is able to effectively resolve alignments in regions containing nearby homologous elements-- a particularly challenging problem in read mapping-- through the introduction of a novel statistical binning optimization framework, which allows us to find variants in the pharmacogenomically important CYP2D region that go undetected when using Lariat or BWA. Lastly, we show that EMAs alignments improve phasing performance compared to Lariats in both NA12878 and NA24385, producing fewer switch/mismatch errors and larger phase blocks on average.\n\nEMA software and datasets used are available at http://ema.csail.mit.edu.

bioinformatics

Metagenomic binning through low density hashing

Bacterial microbiomes of incredible complexity are found throughout the world, from exotic marine locations to the soil in our yards to within our very guts. With recent advances in Next-Generation Sequencing (NGS) technologies, we have vastly greater quantities of microbial genome data, but the nature of environmental samples is such that DNA from different species are mixed together. Here, we present Opal for metagenomic binning, the task of identifying the origin species of DNA sequencing reads. Our Opal method introduces low-density, even-coverage hashing to bioinformatics applications, enabling quick and accurate metagenomic binning. Our tool is up to two orders of magnitude faster than leading alignment-based methods at similar or improved accuracy, allowing computational tractability on large metagenomic datasets. Moreover, on public benchmarks, Opal is substantially more accurate than both alignment-based and alignment-free methods (e.g. on SimHC20.500, Opal achieves 95% F1-score while Kraken and CLARK achieve just 91% and 88%, respectively); this improvement is likely due to the fact that the latter methods cannot handle computationally-costly long-range dependencies, which our even-coverage, low-density fingerprints resolve. Notably, capturing these long-range dependencies drastically improves Opals ability to detect unknown species that share a genus or phylum with known bacteria. Additionally, the family of hash functions Opal uses can be generalized to other sequence analysis tasks that rely on k-mer based methods to encode long-range dependencies.

bioinformatics

C. elegans exhibits coordinated oscillation in gene expression during development

BackgroundThe advent of in vivo automated single-cell lineaging and sequencing will dramatically increase our understanding of development. New integrative analysis techniques are needed to generate insights from single-cell developmental data.\n\nResultsWe applied novel meta-analysis techniques to the EPIC single-cell-resolution developmental gene expression dataset for C. elegans to show that a simple linear combination of the expression levels of the developmental genes is strongly correlated with the developmental age of the organism, irrespective of the cell division rate of different cell lineages. We uncovered a pattern of collective sinusoidal oscillation in gene activation, in multiple dominant frequencies and in multiple orthogonal axes of gene expression, pointing to the existence of a coordinated, multi-frequency global timing mechanism. We developed a novel method based on Fishers Discriminant Analysis (FDA) to identify linear gene expression weightings that are able to produce sinusoidal oscillations of any frequency and phase, adding to the evidence that oscillatory mechanisms likely play an important role in the timing of development. We cross-linked EPIC with gene ontology and anatomy ontology terms, employing FDA methods to identify previously unknown positive and negative genetic contributions to developmental processes and cell phenotypes.\n\nConclusionsThis meta-analysis demonstrates new evidence for direct linear and/or sinusoidal mechanisms regulating the timing of development. We uncovered a number of previously unknown positive and negative correlations between developmental genes and developmental processes or cell phenotypes. The presented novel analysis techniques are broadly applicable within developmental biology.

developmental biology