bioRxiv Science⌕ Search

Biology subjects

Chang, K.-L.

Publications and source records attributed to Chang, K.-L..

2 recordsLinked to original sources

Vision-Based Genomic Model for Copy Number Variant Pathogenicity Prediction

Copy number variants (CNVs) are a major class of structural genomic alterations underlying rare disease, including neurodevelopmental delay and intellectual disability, yet predicting their pathogenicity remains challenging. Existing methods reduce CNVs to region-level numerical features, discarding the positional structure and cross-track patterns that expert clinical reviewers use to interpret genomic evidence. To address this, we introduce TO_SCPLOWESSERACTC_SCPLOW for CNV, a track-based spatial representation for CNV pathogenicity prediction, which represents each variant as a base-pair-resolution multi-track image and models spatial genomic patterns across annotation tracks while preserving positional structure and cross-track dependencies. Trained on a chromosome-level hold-out split of the ClinVar dataset, TO_SCPLOWESSERACTC_SCPLOW outperforms prior methods on held-out and curated noncoding benchmarks, improving AUROC by up to 0.10 over the state-of-the-art baseline. On the independent DECIPHER cohort, the model demonstrates generalizability by maintaining the highest AUROC and the highest F1 score across baselines. Furthermore, our model localizes pathogenic signals to clinically meaningful genomic subregions, providing track-annotated evidence that supports practical clinical interpretation.

bioinformatics↗

Reinforcement learning enables single-cell foundation models to learn cellular differentiation

While single-cell foundation models excel at static representation learning and single-step perturbation prediction, their capacity to model and control dynamic, sequential cell state transitions remains underexplored. Here, we introduce Differentiation with Reinforcement Learning (DiRL), a framework that transforms foundation models into predictive environments for optimizing multi-step perturbation strategies. By formulating differentiation as a goal-conditioned sequential decision-making problem, DiRL trains agents to navigate the high-dimensional latent landscape of gene expression, learning policies that direct stem cells toward specific terminal fates. Evaluated on iPSC-derived organoid datasets, DiRL outperforms random perturbation baselines with a 69% win rate, with performance scaling systematically as the planning horizon increases. Beyond optimization, DiRL offers interpretable insights into the mechanics of cell fate transitions: the learned policies successfully recover known sequential regulators in Wnt and Hedgehog signaling pathways, while the value function from the critic model recapitulates biological pseudotime orderings comparable to established trajectory inference methods. These results demonstrate that coupling reinforcement learning with foundation models enables a paradigm shift from static embedding analysis to dynamic trajectory modeling, providing a powerful engine for discovering sequential perturbation strategies in regenerative medicine.

bioinformatics↗