bioRxiv Science⌕ Search

Biology subjects

Buralkin, I.

Publications and source records attributed to Buralkin, I..

2 recordsLinked to original sources

Vision-Based Genomic Model for Copy Number Variant Pathogenicity Prediction

Copy number variants (CNVs) are a major class of structural genomic alterations underlying rare disease, including neurodevelopmental delay and intellectual disability, yet predicting their pathogenicity remains challenging. Existing methods reduce CNVs to region-level numerical features, discarding the positional structure and cross-track patterns that expert clinical reviewers use to interpret genomic evidence. To address this, we introduce TO_SCPLOWESSERACTC_SCPLOW for CNV, a track-based spatial representation for CNV pathogenicity prediction, which represents each variant as a base-pair-resolution multi-track image and models spatial genomic patterns across annotation tracks while preserving positional structure and cross-track dependencies. Trained on a chromosome-level hold-out split of the ClinVar dataset, TO_SCPLOWESSERACTC_SCPLOW outperforms prior methods on held-out and curated noncoding benchmarks, improving AUROC by up to 0.10 over the state-of-the-art baseline. On the independent DECIPHER cohort, the model demonstrates generalizability by maintaining the highest AUROC and the highest F1 score across baselines. Furthermore, our model localizes pathogenic signals to clinically meaningful genomic subregions, providing track-annotated evidence that supports practical clinical interpretation.

bioinformatics↗

scDeepVariant: A population-informed deep learning framework for germline variant calling in scRNA-seq

Single-cell RNA sequencing (scRNA-seq) provides unprecedented resolution of cellular heterogeneity while also capturing information on germline genetic variation, but accurate variant calling remains limited by sparse coverage, allelic imbalance, and RNA-specific artifacts. Existing single-cell methods, including cellSNP, scAllele, and Monopogen, address some of these challenges, yet either suffer from low sensitivity and precision or rely on linkage disequilibrium (LD) priors that restrict performance on rare variants. Here, we introduce scDeepVariant (scDV), a deep learning-based framework adapted from DeepVariant and trained on paired whole-genome sequencing (WGS) and single-nucleus RNA sequencing (snRNA-seq) data. We show that scDV can be effectively trained on sparse single-cell data and that augmenting models with allele frequency information from gnomAD or the 1000 Genomes Project consistently improves performance. Across benchmarks, scDV with allele frequency channels achieved higher precision and recall than standard six-channel configurations, surpassing Monopogen at coverage depths above 10x and demonstrating a pronounced advantage in rare variant detection, where LD-based refinement is most limited. These results establish scDV as a robust alternative for germline variant discovery from scRNA-seq and highlight the broader value of integrating population-scale information into deep learning frameworks for transcriptomic variant calling.

bioinformatics↗