bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.09.03.674100

DAESC+: High-performance, integrated software for single-cell allele-specific expression data

Abstract

Single-cell allele-specific expression (ASE) provides valuable insights into gene regulatory mechanisms. However, its utility is limited by the lack of dedicated computational tools. We present DAESC+, an end-to-end software package for the processing and analysis of single-cell ASE. The preprocessing module, DAESC-P, is a user-friendly bioinformatics pipeline to obtain ASE counts from multiplexed scRNA-seq data. The analysis module, DAESC-GPU, is a scalable tool for differential ASE analysis powered by graphics processing units (GPUs). We demonstrated that DAESC-P is more accurate than the existing SALSA pipeline. DAESC-GPU is dozens of times faster than our previous method (DAESC) and scalable to over a million cells. Applying DAESC+ to a subset of the OneK1K cohort, we identified 15 genes exhibiting differential regulatory patterns between naive and central memory CD4+ T cells, and 2 genes between naive and memory B cells.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cui, T., Qi, G.. 2025-09-09. DAESC+: High-performance, integrated software for single-cell allele-specific expression data. https://doi.org/10.1101/2025.09.03.674100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Single-Sample Network Topology Unveils Pathogenic Hubs in Sjögren's Disease

Sjogren's disease (SD) is a systemic autoimmune condition characterized by extensive clinical and biological heterogeneity, complicating the development of targeted therapies. To better describe the personalized molecular rewiring driving SD pathogenesis, we utilized bulk RNA-sequencing data from whole blood samples of the PRECISESADS IMI consortium to construct sample-specific gene interaction networks by applying the LIONESS algorithm integrated with experimental protein-protein interaction evidence. Comparing SD patients against healthy controls, we identified an absolute total of 647 Differentially Interacting Genes (DIGs) representing structural shifts in network connectivity. Topological clustering of these DIGs revealed four major functional macro-modules underlying disease pathology: defense response to virus, B cell activation, DNA replication, and positive regulation of protein catabolic process. Central topological hubs, such as ISG15, OASL, UBE2L6, and LGALS3BP, were identified as drivers of this reorganization, a fact supported by integrating network topology with single-sample Gene Set Enrichment Analysis (ssGSEA), which demonstrated that the interaction degree of these hubs exhibits strong positive correlations with the functional activity of their respective pathogenic pathways. To translate these structural findings into therapeutic opportunities, we performed an in-silico network vulnerability analysis on patient-specific interaction networks, ranking targets by the fractional loss of global efficiency following their removal and correcting for node degree. This approach prioritised the receptor tyrosine kinase EPHB2, which has no prior description in this disease, alongside the kinase BLK and the B-cell co-receptor CD22, a target independently evaluated in a previous clinical trial in SD. Thus, our single-sample network framework provides a computational mechanism to uncover pathogenesis-related modules and to prioritise drug repurposing candidates in SD, while also recovering targets whose clinical failure bounds what network topology alone can predict.

bioinformatics↗

EPIC: An open community challenge for sequence-based prediction of transcription initiation in five non-model metazoans

Predicting gene expression from DNA sequence is a central problem with critical biological and clinical implications. Recent sequence-to-function models are reported to achieve improved performance, yet it remains unclear how much of what they learn reflects genuine principles of eukaryotic transcription initiation and to what extent they are able to generalize beyond humans and other primary model species. A fair and blind benchmark has likewise been missing. Here we introduce EPIC, the Eukaryotic Promoter and transcription Initiation prediction Challenge. Teams receive strand-specific, single-nucleotide-resolution initiation profiles for 80-95% of the genome and are asked to predict the remainder from DNA sequence alone. To level the field, the challenge relies on understudied animals spanning three phyla: octopus, oyster, milkweed bug, Indian meal moth, and shark. EPIC is open to everyone and closes on December 31, 2026. All teams clearing the dinucleotide precision baseline are invited to join the consortium authorship of the post-challenge publication, and winners are invited for personal authorship.

bioinformatics↗

Kryptix-1: Conformational Ensemble Sampling Substantially Improves Cryptic Pocket Detection Relative to Static-Structure and Prior Computational Approaches

Cryptic binding pockets; sites absent or occluded in a proteins resting-state structure that become druggable only in specific, transiently populated conformations; represent one of the largest untapped opportunities in structure-based drug discovery. Large-scale structural surveys estimate that cryptic pockets occur in on the order of one in six protein families genome-wide, and that accounting for them could expand the druggable fraction of the disease-associated human proteome from roughly 40% to nearly 80%. Detecting cryptic pockets computationally has historically forced a choice btw physics-based conformational sampling, which is reliable but too computationally expensive to deploy across more than a handful of targets, and fast machine-learning pocket predictors trained on static structures, which frequently over-predict and lose the precision needed for practical triage. We report Kryptix-1, an in-house pipeline developed at Covenant Biosciences that combines conformational ensemble generation with a consensus pocket-scoring algorithm, benchmarked here against a matched static-structure baseline and against the published literature on a 10-protein sample from CryptoBench, an independently curated cryptic-pocket reference dataset. On the fields own standard residue-overlap metric (Jaccard index 0.5), Kryptix-1 succeeds on 80% of benchmark proteins, compared to 20% for the static baseline and approximately 40-45% for the best-performing methods reported in an independent comparative evaluation using the same metric. We report these results alongside an explicit accounting of where Kryptix-1s advantage is smaller or reverses, and we are explicit throughout that this is a small pilot evaluation, not a fully powered validation study.

bioinformatics↗