bioRxiv Science⌕ Search

Biology subjects

AlKhafaji, A. M.

Publications and source records attributed to AlKhafaji, A. M..

5 recordsLinked to original sources

Scalable single-cell isoform profiling with sequencing-by-expansion

Single-cell RNA sequencing has transformed our understanding of cellular systems, yet the reliance on short-read sequencing restricts analysis to gene-level quantification and obscures the immense biological diversity generated by alternative splicing. While long-read sequencing technologies can capture full-length RNA and resolve transcript isoforms, current platforms remain constrained by throughput and high per-base costs, rendering them impractical for modern million-cell applications. To address this critical limitation, we developed and optimized sequencing-by-expansion (SBX) chemistry for high-throughput single-cell RNA isoform profiling. Integrated within the AXELIOS 1 sequencing platform, SBX employs a unique biochemical conversion process that transforms complementary DNA into expanded surrogate high signal-to-noise polymers called Xpandomers which are sequenced via translocation through a dense nanopore array yielding over 9.5 billion reads in a two-hour run. To leverage this unique data type for long-read single-cell RNA isoform sequencing, we developed the Consensus UMI Deduplication using Longest Length (CUDLL) algorithm, which computationally consolidates variable-length raw SBX reads into single, high-fidelity consensus reads, elevating sequence accuracy to 99.83% and maximizing per transcript read length. We demonstrate that this consensus approach successfully captures the vast isoform diversity of single-cell libraries and enables the accurate measure of differential isoform expression across distinct cell types in peripheral blood mononuclear cells. Furthermore, SBX coupled with CUDLL efficiently resolves T-cell and B-cell receptor clonotypes directly from whole-transcriptome libraries without the need for VDJ-specific target enrichment. Ultimately, this work establishes SBX and the AXELIOS 1 as a transformative platform for high-scale single-cell isoform sequencing.

genetics↗

Stitch-seq: Scalable CRISPR gene expression response profiling

Single-cell profiling of genetic perturbations has expanded our ability to map causal links between genes and phenotypes; however, the high cost and technical complexity of current methods restrict systematic interrogation of dynamic cellular programs. Here, we present Stitch-seq, a high-throughput pooled functional genomics sequencing method enabling simultaneous capture of CRISPR perturbations and targeted gene and protein expression across millions of cells. Stitch-seq utilizes single-cell droplet-based overlap-extension reverse-transcription PCR reactions to physically link gene expression features of interest to perturbation identifiers without cell barcoding or extensive sequencing. We validated Stitch-seqs high fidelity using simplified models, benchmarked multi-omic Stitch-seq against single-cell RNA-sequencing in the MCF10A Epithelial-Mesenchymal Transition (EMT) model, and applied Stitch-seq to map transcriptional responses of MCF10A cells undergoing TGF-{beta}-induced EMT to perturbations across five time points. By efficiently delivering large-scale multi-omic gene expression readouts, Stitch-seq provides a powerful and accessible modality for the routine dissection of complex biological pathways.

bioengineering↗

Accurate strand-specific long-read transcript isoform discovery and quantification at bulk, single-cell, and single-nucleus resolution

Recent advances in long-read transcriptome sequencing enable high-throughput profiling of full-length RNA isoforms in bulk, single-cell, and single-nucleus samples. However, long-read datasets typically contain a mixture of complete and partial transcripts, leading to pervasive ambiguity in read-to-isoform assignment and complicating accurate isoform identification and quantification, particularly in the absence of reliable reference annotations. These challenges are further amplified in single-cell and single-nucleus samples, where coverage is sparse and transcriptional heterogeneity is high. Here, we present the Long Read Alignment Assembler (LRAA), a unified and versatile computational framework for isoform identification and quantification from long-read RNA sequencing data across bulk, single-cell, and single-nucleus transcriptomic samples. LRAA combines splice-graph based structural modeling with expectation maximization based optimization to probabilistically resolve ambiguous read assignments and improve isoform abundance estimation. The framework supports quantification-only, reference-guided, and fully reference-free (de novo) modes of analysis within a single methodological paradigm. We benchmarked LRAA using both simulated and genuine long-read datasets spanning sequencing standards and whole transcriptomes. Central to this evaluation is a novel benchmarking strategy based on Multiplexed Overexpression of Regulatory Factors (MORFs), which provides biologically expressed, barcoded isoforms with unambiguous read-level ground truth. Across all benchmarks, including MORFs, synthetic spike-ins, and whole-transcriptome datasets, LRAA consistently outperformed state-of-the-art methods in isoform identification accuracy, sensitivity, and expression quantification. Finally, we demonstrate the biological utility of LRAA by resolving cell-type-specific isoform usage across peripheral blood immune cell populations and by detecting a pathogenic cryptic isoform of STMN2 with associated transcriptional changes in single-nucleus RNA-seq data from frontal cortex tissue of an individual with frontotemporal dementia (FTD). Together, these results establish LRAA as a robust and general solution for resolving transcript diversity in complex biological systems, from development to disease.

bioinformatics↗

RNA splicing dynamics in CD8 T cells uncovers isoforms that impact T cell-mediated cancer immunotherapy

Immune checkpoint blockade has transformed cancer therapy, yet many patients fail to respond, and few new targets have emerged beyond PD-1 and CTLA-4. Alternative splicing dramatically diversifies the T cell proteome, but the functional roles of most isoforms remain unknown. Here we constructed the first single-cell splicing atlas of human CD8 T cells, capturing dynamic isoform programs across activation and subset differentiation. This revealed distinct splicing footprints that refine conventional transcriptomic states and highlight receptor families with isoform-level regulation. To functionally interrogate these candidates, we developed SpliceSeek, a CRISPR-based pooled screening platform that perturbs splice sites to redirect isoform usage. Using SpliceSeek, we uncovered isoform-specific immune checkpoints whose perturbation enhanced effector function and tumor control, including the LRRN3-203 isoform, which augmented cytokine secretion and antitumor immunity in mice models. Together, our results establish alternative splicing as a targetable layer of immune regulation and demonstrate the potential of isoform-focused screening to expand the landscape of cancer immunotherapy.

immunology↗

CTAT-LR-fusion: accurate fusion transcript identification from long and short read isoform sequencing at bulk or single cell resolution

Gene fusions are found as cancer drivers in diverse adult and pediatric cancers. Accurate detection of fusion transcripts is essential in cancer clinical diagnostics, prognostics, and for guiding therapeutic development. Most currently available methods for fusion transcript detection are compatible with Illumina RNA-seq involving highly accurate short read sequences. Recent advances in long read isoform sequencing enable the detection of fusion transcripts at unprecedented resolution in bulk and single cell samples. Here we developed a new computational tool CTAT-LR-fusion to detect fusion transcripts from long read RNA-seq with or without companion short reads, with applications to bulk or single cell transcriptomes. We demonstrate that CTAT-LR-fusion exceeds fusion detection accuracy of alternative methods as benchmarked with simulated and real long read RNA-seq. Using short and long read RNA-seq, we further apply CTAT-LR-fusion to bulk transcriptomes of nine tumor cell lines, and to tumor single cells derived from a melanoma sample and three metastatic high grade serous ovarian carcinoma samples. In both bulk and in single cell RNA-seq, long isoform reads yielded higher sensitivity for fusion detection than short reads with notable exceptions. By combining short and long reads in CTAT-LR-fusion, we are able to further maximize detection of fusion splicing isoforms and fusion-expressing tumor cells. CTAT-LR-fusion is available at https://github.com/TrinityCTAT/CTAT-LR-fusion/wiki.

cancer biology↗