bioRxiv Science⌕ Search

Biology subjects

Farooqi, M. S.

Publications and source records attributed to Farooqi, M. S..

4 recordsLinked to original sources

Hunting for microsatellite instability in long-read data with Owl

Microsatellite instability (MSI) is a key biomarker of mismatch repair deficiency and response to immunotherapy, yet most existing genomic detection methods are optimized for short-read sequencing and rely on small panels of homopolymer markers, limiting the ability to characterize genome-wide and motif-specific patterns of instability. Here we present Owl, a bioinformatic tool for quantifying MSI from long-read (PacBio) genomic data. Owl leverages a genome-wide marker set of more than 140,000 microsatellite repeats ranging from 1-6 bp in length to measure MSI across a phased genome. Using a wrap-around alignment algorithm, Owl constructs repeat-length distributions at each marker site and flags somatic instability using the coefficient of variation. We applied Owl to screen for markers with stable coverage, phasing, and baseline variation across 131 diverse genomes from the Human Pangenome Reference Consortium, where Owl scores ranged from 1.4% to 5.4% of markers exceeding the instability threshold. When applied to 19 cancer cell lines and one diffuse astrocytoma tumor-normal pair, Owl identified five MSI-high genomes with 15-18% unstable markers and showed close concordance with an Illumina DRAGEN MSI assay for the astrocytoma sample. Motif-level analyses revealed shared enrichment of short homopolymer and dinucleotide (A- and AT-rich) repeats across MSI-high cancers, and additionally uncovered a distinct pattern of elevated GGAA microsatellite instability in Ewing sarcoma cell lines, consistent with the known role of the EWS::FLI1 fusion protein at GGAA-rich regulatory elements. Owl is implemented in Rust and integrated into the PacBio HiFi Somatic workflow, providing a scalable framework for MSI analysis from long-read sequencing focused on repeat instability specifically in tumor samples.

bioinformatics↗

A comprehensive bulk and single-cell transcriptional atlas of pediatric leukemias

Advancements in understanding the molecular factors driving pediatric leukemias have led to an ever-increasing volume and diversity of data being generated. Recent studies are moving beyond DNA-based profiling to incorporate transcriptional data, enhancing the characterization of these cancers and informing clinical decisions. However, many existing datasets focus on specific leukemia subtypes, limiting the extraction of broader clinically relevant insights. Here, we present a comprehensive dataset of bulk and single-cell transcriptional data from 69 pediatric leukemia patients, encompassing eight leukemias. This is one of the most diverse pediatric leukemia datasets published to date. Through comparative analysis and explainable machine learning, we demonstrate the utility of these datasets in improving pediatric leukemia characterization and supporting the integration of transcriptional data into clinical testing.

cancer biology↗

Integrated multi-omic analysis reveals novel subtype-specific regulatory interactions in pediatric B-cell acute lymphoblastic leukemia

Molecular subtyping of pediatric B-cell acute lymphoblastic leukemia (B-ALL) has improved patient outcomes through stratification and selection of targeted therapies. Despite extensive genomic and transcriptomic profiling of this cancer, few studies to date have characterized the proteomic landscape, although proteins are the direct targets of many therapeutic agents. In this study, we demonstrate the utility of multi-omic integration of global transcriptomic, proteomic, and phosphoproteomic profiles of samples from patients diagnosed with either of two B-ALL subtypes - Ph- like (BCR::ABL1-like) and ETV6::RUNX1. Through individual and multi-omic analysis, we recapitulate known transcriptomic findings and identify novel subtype-specific proteomic and phosphoproteomic biomarkers. Our findings suggest a previously undescribed role for calcium-dependent signaling processes in Ph-like B-ALL, which has the potential to serve as a novel avenue for targeted treatments. By integrating multiple omics modalities, we identify not only features of interest but also begin to unravel the regulatory interactions driving subtype-specific mechanisms of leukemogenesis. This integrated analytic approach paves the way for enhanced precision medicine for precise subtyping and treatment selection for pediatric leukemia patients. Mass spectrometry data generated in this study have been deposited in MassIVE under accession MSV000097955.

cancer biology↗

DeepSomatic: Accurate somatic small variant discovery for multiple sequencing technologies

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies now offer potential advantages in terms of repeat mapping and variant phasing. We present DeepSomatic, a deep learning method for detecting somatic SNVs and insertions and deletions (indels) from both short-read and long-read data, with modes for whole-genome and exome sequencing, and able to run on tumor-normal, tumor-only, and with FFPE-prepared samples. To help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available a dataset of five matched tumor-normal cell line pairs sequenced with Illumina, PacBio HiFi, and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples and technologies (short-read and long-read), DeepSomatic consistently outperforms existing callers, particularly for indels.

bioinformatics↗