bioRxiv Science⌕ Search

Biology subjects

Baharav, T. Z.

Publications and source records attributed to Baharav, T. Z..

5 recordsLinked to original sources

SPLASH: a statistical, reference-free genomic algorithm unifies biological discovery

The authors have withdrawn this manuscript due to a duplicate posting of manuscript number BIORXIV/2022/497555. Therefore, the authors do not wish this work to be cited as reference for the project. If you have any questions, please contact the corresponding author. The correct preprint can be found at doi: https://doi.org/10.1101/2022.06.24.497555

genomics↗

NOMAD2 provides ultra-efficient, scalable, and unsupervised discovery on raw sequencing reads

SPLASH is an unsupervised, reference-free, and unifying algorithm that discovers regulated sequence variation through statistical analysis of k-mer composition, subsuming many application-specific methods. Here, we introduce SPLASH2, a fast, scalable implementation of SPLASH based on an efficient k-mer counting approach. SPLASH2 enables rapid analysis of massive datasets from a wide range of sequencing technologies and biological contexts, delivering unparalleled scale and speed. The SPLASH2 algorithm unveils new biology (without tuning) in single-cell RNA-sequencing data from human muscle cells, as well as bulk RNA-seq from the entire Cancer Cell Line Encyclopedia (CCLE), including substantial unannotated alternative splicing in cancer transcriptome. The same untuned SPLASH2 algorithm recovers the BCR-ABL gene fusion, and detects circRNA sensitively and specifically, underscoring SPLASH2s unmatched precision and scalability across diverse RNA-seq detection tasks.

bioinformatics↗

An interpretable, finite sample valid alternative to Pearson's X2 for scientific discovery

Contingency tables, data represented as counts matrices, are ubiquitous across quantitative research and data-science applications. Existing statistical tests are insufficient however, as none are simultaneously computationally efficient and statistically valid for a finite number of observations. In this work, motivated by a recent application in reference-free genomic inference (1), we develop OASIS (Optimized Adaptive Statistic for Inferring Structure), a family of statistical tests for contingency tables. OASIS constructs a test-statistic which is linear in the normalized data matrix, providing closed form p-value bounds through classical concentration inequalities. In the process, OASIS provides a decomposition of the table, lending interpretability to its rejection of the null. We derive the asymptotic distribution of the OASIS test statistic, showing that these finitesample bounds correctly characterize the test statistics p-value up to a variance term. Experiments on genomic sequencing data highlight the power and interpretability of OASIS. The same method based on OASIS significance calls detects SARS-CoV-2 and Mycobacterium Tuberculosis strains de novo, which cannot be achieved with current approaches. We demonstrate in simulations that OASIS is robust to overdispersion, a common feature in genomic data like single cell RNA-sequencing, where under accepted noise models OASIS still provides good control of the false discovery rate, while Pearsons X2 test consistently rejects the null. Additionally, we show on synthetic data that OASIS is more powerful than Pearsons X2 test in certain regimes, including for some important two group alternatives, which we corroborate with approximate power calculations. Significance StatementContingency tables are pervasive across quantitative research and data-science applications. Existing statistical tests fall short, however; none provide robust, computationally efficient inference and control Type I error. In this work, motivated by a recent advance in reference-free inference for genomics, we propose a family of tests on contingency tables called OASIS. OASIS utilizes a linear test-statistic, enabling the computation of closed form p-value bounds, as well as a standard asymptotic normality result. OASIS provides a partitioning of the table for rejected hypotheses, lending interpretability to its rejection of the null. In genomic applications, OASIS performs reference-free and metadata-free variant detection in SARS-CoV-2 and M. Tuberculosis, and demonstrates robust performance for single cell RNA-sequencing, all tasks without existing solutions.

bioinformatics↗

Unsupervised reference-free inference reveals unrecognized regulated transcriptomic complexity in human single cells

Myriad mechanisms diversify the sequence content of eukaryotic transcripts at both the DNA and RNA levels, leading to profound functional consequences. Examples of this diversity include RNA splicing and V(D)J recombination. Currently, these mechanisms are detected using fragmented bioinformatic tools that require predefining a form of transcript diversification and rely on alignment to an incomplete reference genome, filtering out unaligned sequences, potentially crucial for novel discoveries. Here, we present SPLASH+, significantly advancing biological discovery possible with SPLASH, our recently introduced efficient, reference-free statistical approach. Integrating a micro-assembly and biological interpretation framework, SPLASH+ enables new discoveries including broad and novel examples of transcript diversification in single cells de novo, without the need for cell type metadata, which is impossible with current algorithms. Applied to 10,326 primary human single cells across 19 tissues profiled with SmartSeq2, SPLASH+ discovers a set of splicing and histone regulators with highly conserved intronic regions that are themselves subject to complex splicing regulation. Additionally, it reveals unreported transcript diversity in the heat shock protein HSP90AA1, as well as diversification in centromeric RNA expression, V(D)J recombination, RNA editing, and repeat expansion, all missed by existing methods. SPLASH+ is highly efficient, enabling the discovery of an unprecedented breadth of RNA regulation and diversification in single cells through a new automated paradigm of unbiased transcriptomic analysis.

bioinformatics↗

A statistical, reference-free algorithm subsumes myriad problems in genome science and enables novel discovery

Todays genomics workflows typically require alignment to a reference sequence, which limits discovery. We introduce a new unifying paradigm, SPLASH (Statistically Primary aLignment Agnostic Sequence Homing), an approach that directly analyzes raw sequencing data to detect a signature of regulation: sample-specific sequence variation. The approach, which includes a new statistical test, is computationally efficient and can be run at scale. SPLASH unifies detection of myriad forms of sequence variation. We demonstrate that SPLASH identifies complex mutation patterns in SARS-CoV-2 strains, discovers regulated RNA isoforms at the single cell level, documents the vast sequence diversity of adaptive immune receptors, and uncovers biology in non-model organisms undocumented in their reference genomes: geographic and seasonal variation and diatom association in eelgrass, an oceanic plant impacted by climate change, and tissue-specific transcripts in octopus. SPLASH is a new unifying approach to genomic analysis that enables an expansive scope of discovery without metadata or references. One-sentence summarySPLASH is a unifying, statistically driven approach to biological discovery from raw sequencing data, bypassing alignment.

genomics↗