bioRxiv ScienceSearch

Biology subjects

Batzoglou, S.

Publications and source records attributed to Batzoglou, S..

8 recordsLinked to original sources

Network Enhancement: a general method to denoise weighted biological networks

Networks are ubiquitous in biology where they encode connectivity patterns at all scales of organization, from molecular to the biome. However, biological networks are noisy due to the limitations of technology used to generate them as well as inherent variation within samples. The presence of high levels of noise can hamper discovery of patterns and dynamics encapsulated by these networks. Here we propose Network Enhancement (NE), a novel method for improving the signal-to-noise ratio of undirected, weighted networks, and thereby improving the performance of downstream analysis. NE applies a novel operator that induces sparsity and leverages higher-order network structures to remove weak edges and enhance real connections. This iterative approach has a closed-form solution at convergence with desirable performance properties. We demonstrate the effectiveness of NE in denoising biological networks for several challenging yet important problems. Our experiments show that NE improves gene function prediction by denoising interaction networks from 22 human tissues. Further, we use NE to interpret noisy Hi-C contact maps from the human genome and demonstrate its utility across varying degrees of data quality. Finally, when applied to fine-grained species identification, NE outperforms alternative approaches by a significant margin. Taken together, our results indicate that NE is widely applicable for denoising weighted biological networks, especially when they contain high levels of noise.

bioinformatics

Multi-omic tumor data reveal diversity of molecular mechanisms underlying survival

Outcomes for cancer patients vary greatly even within the same tumor type, and characterization of molecular subtypes of cancer holds important promise for improving prognosis and personalized treatment. This promise has motivated recent efforts to produce large amounts of multidimensional genomic ( multi-omic) data, but current algorithms still face challenges in the integrated analysis of such data. Here we present Cancer Integration via Multikernel Learning (CIMLR), a new cancer subtyping method that integrates multi-omic data to reveal molecular subtypes of cancer. We apply CIMLR to multi-omic data from 36 cancer types and show significant improvements in both computational efficiency and ability to extract biologically meaningful cancer subtypes. The discovered subtypes exhibit significant differences in patient survival for 27 of 36 cancer types. Our analysis reveals integrated patterns of gene expression, methylation, point mutations and copy number changes in multiple cancers and highlights patterns specifically associated with poor patient outcomes.

cancer biology

Culture-free generation of microbial genomes from human and marine microbiomes

Our understanding of natural microbial communities is shaped by the careful investigation of a relatively small number of isolated and cultured organisms, and by analysis of genomic sequences obtained by culture-free metagenomic sequencing approaches. Metagenomic shotgun sequencing has facilitated partial reconstruction of strain-level community structure and functional repertoire. Unfortunately, it remains difficult to cost-effectively produce high quality genome drafts for individual microbes without isolation and culture. Recent molecular techniques that partition long DNA fragments and then barcode short fragments derived from them produce \"read clouds\", which are short-read sequences containing long-range information. Here, we present a novel application of a read cloud technique to microbiome samples, as well as Athena, a de novo assembler that uses these barcodes to produce improved metagenomic assemblies. We apply our approach to sequence human stool samples from two healthy individuals, and compare it to existing short read and synthetic long read metagenomic sequencing approaches. We find that read cloud metagenomic sequencing and Athena assembly produce the most complete individual genome drafts. These genome drafts are also highly contiguous (>200kb N50, <10 contigs), even for bacteria that have relatively low (20x) raw short read sequence coverage. We also apply this approach to a significantly more complex marine sediment sample and obtain 23 genome drafts with valuable 16S ribosomal RNA taxonomic marker sequences, nine of which are complete genome drafts. Read cloud metagenomic sequencing allows culture-free generation of high quality microbial genome drafts using only a single shotgun experiment.

genomics

HAPDeNovo: a haplotype-based approach for filtering and phasing de novo mutations in linked read sequencing data

BackgroundDe novo mutations (DNMs) are associated with neurodevelopmental and congenital diseases, and their detection can contribute to understanding disease pathogenicity. However, accurate detection is challenging because of their small number relative to the genome-wide false positives in next generation sequencing (NGS) data. Software such as DeNovoGear and TrioDeNovo have been developed to detect DNMs, but at good sensitivity they still produce many false positive calls.\n\nResultsTo address this challenge, we develop HAPDeNovo, a program that leverages phasing information from linked read sequencing, to remove false positive DNMs from candidate lists generated by DNM-detection tools. Short reads from each phasing block are allocated to each of the two haplotypes followed by generating a haploid genotype for each putative DNM.HAPDeNovo removes variants that are called as heterozygous in one of the haplotypes because they are almost certainly false positives. Our experiments on 10X Chromium linked read sequencing trio data reveal that HAPDeNovo eliminates 80% to 99% of false positives regardless of how large the candidate DNM set is.\n\nConclusionsHAPDeNovo leverages the haplotype information from linked read sequencing to remove spurious false positive DNMs effectively, and it increases accuracy of DNM detection dramatically without sacrificing sensitivity.

genomics

GATTACA: Lightweight Metagenomic Binning With Compact Indexing Of Kmer Counts And MinHash-based Panel Selection

We introduce GATTACA, a framework for rapid and accurate binning of metagenomic contigs from a single or multiple metagenomic samples into clusters associated with individual species. The clusters are computed using co-abundance profiles within a set of reference metagnomes; unlike previous methods, GATTACA estimates these profiles from k-mer counts stored in a highly compact index. On multiple synthetic and real benchmark datasets, GATTACA produces clusters that correspond to distinct bacterial species with an accuracy that matches earlier methods, while being up to 20x faster when the reference panel index can be computed offline and 6x faster for online co-abundance estimation. Leveraging the MinHash technique to quickly compare metagenomic samples, GATTACA also provides an efficient way to identify publicly-available metagenomic data that can be incorporated into the set of reference metagenomes to further improve binning accuracy. Thus, enabling easy indexing and reuse of publicly-available metagenomic datasets, GATTACA makes accurate metagenomic analyses accessible to a much wider range of researchers.

bioinformatics

De novo assembly of microbial genomes from human gut metagenomes using barcoded short read sequences

Although shotgun short-read sequencing has facilitated the study of strain-level architecture within complex microbial communities, existing metagenomic approaches often cannot capture structural differences between closely related co-occurring strains. Recent methods, which employ read cloud sequencing and specialized assembly techniques, provide significantly improved genome drafts and show potential to capture these strain-level differences. Here, we apply this read cloud metagenomic approach to longitudinal stool samples from a patient undergoing hematopoietic cell transplantation. The patients microbiome is profoundly disrupted and is eventually dominated by Bacteroides caccae. Comparative analysis of B. caccae genomes obtained using read cloud sequencing together with metagenomic RNA sequencing allows us to predict that particular mobile element integrations result in increased antibiotic resistance, which we further support using in vitro antibiotic susceptibility testing. Thus, we find read cloud sequencing to be useful in identifying strain-level differences that underlie differential fitness.

bioinformatics

SIMLR: A Tool For Large-Scale Single-Cell Analysis By Multi-Kernel Learning

MotivationWe here present SIMLR (Single-cell Interpretation via Multi-kernel LeaRning), an open-source tool that implements a novel framework to learn a cell-to-cell similarity measure from single-cell RNA-seq data. SIMLR can be effectively used to perform tasks such as dimension reduction, clustering, and visualization of heterogeneous populations of cells. SIMLR was benchmarked against state-of-the-art methods for these three tasks on several public datasets, showing it to be scalable and capable of greatly improving clustering performance, as well as providing valuable insights by making the data more interpretable via better a visualization.\n\nAvailability and ImplementationSIMLR is available on GitHub in both R and MATLAB implementations. Furthermore, it is also available as an R package on bioconductor.org.\n\nContactbowang87@stanford.edu or daniele.ramazzotti@stanford.edu\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

Large-scale Population Genotyping from Low-coverage Sequencing Data using a Reference Panel

In recent years, several large-scale whole-genome projects sequencing tens of thousands of individuals were completed, with larger studies are underway. These projects aim to provide high-quality genotypes for a large number of whole genomes in a cost-efficient manner, by sequencing each genome at low coverage and subsequently identifying alleles jointly in the entire cohort. Here we present Ref-Reveel, a novel method for large-scale population genotyping. We show that Ref-Reveel provides genotyping at a higher accuracy and higher efficiency in comparison to existing methods by applying our method to one of the largest whole-genome sequencing datasets presently available to the public. We further show that utilizing the resulting genotype panel as references, through the Ref-Reveel framework, greatly improves the ability to call genotypes accurately on newly sequenced genomes. In addition, we present a Ref-Reveel pipeline that is applicable for genotyping of very small datasets. In summary, Ref-Reveel is an accurate, scalable and applicable method for a wide range of genotyping scenarios, and will greatly improves the quality of calling genomic alterations in current and future large-scale sequencing projects.

bioinformatics