bioRxiv Science⌕ Search

Biology subjects

Ling, M. H.

Publications and source records attributed to Ling, M. H..

2 recordsLinked to original sources

Isoform-level discovery, quantification and fusion analysis from single-cell and spatial long-read RNA-seq data with Bambu-Clump

Single cell and spatial transcriptomics have dramatically changed how we can profile RNA from heterogenous biological samples. Combining single cell and spatial profiling with long read RNA-Seq promises to enable the discovery and quantification of individual RNA isoforms at the single-cell level. However, highly multiplexed data such as from a single cell experiment only generates a limited number of reads for each cell, constituting a major challenge for transcript discovery and quantification with existing approaches that usually have limited power for samples with low sequencing depth. Here we present Bambu-Clump, a computational method that performs transcript discovery and quantification from single cell and spatial long read RNA-Seq data using information from both each cell and the cell cluster. Using this approach, Bambu-Clump provides the most accurate transcript discovery compared to other existing methods, and improves transcript quantification compared to methods that rely on estimates derived from single cells. We apply Bambu-Clump to identify fusion transcripts in single-cells, compare 5 and 3 selection protocols, and identify novel isoform cell-type markers in spatial mouse brain data. Together, Bambu-Clump provides an easy-to-use, efficient, and accurate method for analysing individual isoform expression for single cells and cell clusters across multiple datasets and replicates from long read RNA-Seq.

bioinformatics↗

Context-Aware Transcript Quantification from Long Read RNA-Seq data with Bambu

Most approaches to transcript quantification rely on fixed reference annotations. However, the transcriptome is dynamic, and depending on the context, such static annotations contain inactive isoforms for some genes while they are incomplete for others. To address this, we have developed Bambu, a method that performs machine-learning based transcript discovery to enable quantification specific to the context of interest using long-read RNA-Seq data. To identify novel transcripts, Bambu employs a precision-focused threshold referred to as the novel discovery rate (NDR), which replaces arbitrary per-sample thresholds with a single interpretable parameter. Bambu retains the full-length and unique read counts, enabling accurate quantification in presence of inactive isoforms. Compared to existing methods for transcript discovery, Bambu achieves greater precision without sacrificing sensitivity. We show that context-aware annotations improve abundance estimates for both novel and known transcripts. We apply Bambu to human embryonic stem cells to quantify isoforms from repetitive HERVH-LTR7 retrotransposons, demonstrating the ability to estimate transcript expression specific to the context of interest.

bioinformatics↗