bioRxiv · 10.64898/2026.08.19.745697
A self-supervised DNA foundation model with collapse-resistant multimodal fusion
Abstract
Genomic foundation models pretrained on DNA sequence have achieved strong performance across many tasks, but sequence-only representations cannot fully capture regulatory information from additional DNA-centric modalities. Existing multimodal genomic models are optimized for specific prediction tasks rather than reusable embeddings. Directly fusing heterogeneous modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence embeddings have markedly different statistical structures, making naive alignment prone to near-zero solutions. We present a self-supervised DNA-centric multimodal foundation model integrating DNA sequence embeddings with local and global chromatin accessibility in a shared encoder to produce reusable window-level embeddings. We show that global normalization alleviates this collapse, enabling effective joint learning. The resulting embeddings improve regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline, with further gains on external ClinVar, GTEx eQTL and PBMC caQTL datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, Y.. 2026-08-20. A self-supervised DNA foundation model with collapse-resistant multimodal fusion. https://doi.org/10.64898/2026.08.19.745697
Cite the original work for its findings. Save a collection to share your selection of sources.