bioRxiv · 10.1101/2025.01.06.631595
DNALONGBENCH: A Benchmark Suite for Long-Range DNA Prediction Tasks
Abstract
Modeling long-range DNA dependencies is crucial for understanding genome structure and function across a wide range of biological contexts. However, effectively capturing these extensive dependencies, which may span millions of base pairs in tasks such as three-dimensional (3D) chromatin folding prediction, remains a significant challenge. Furthermore, a comprehensive benchmark suite for evaluating tasks that rely on long-range dependencies is notably absent. To address this gap, we introduce DNALO_SCPLOWONGC_SCPLOWBO_SCPLOWENCHC_SCPLOW, a benchmark dataset encompassing five important genomics tasks that consider long-range dependencies up to 1 million base pairs: enhancer-target gene interaction, expression quantitative trait loci, 3D genome organization, regulatory sequence activity, and transcription initiation signals. To comprehensively assess DNALO_SCPLOWONGC_SCPLOWBO_SCPLOWENCHC_SCPLOW, we evaluate the performance of five methods: a task-specific expert model, a convolutional neural network (CNN)-based model, and three fine-tuned DNA foundation models - HyenaDNA, Caduceus-Ph, and Caduceus-PS. We envision DNALO_SCPLOWONGC_SCPLOWBO_SCPLOWENCHC_SCPLOW as a standardized resource with the potential to facilitate comprehensive comparisons and rigorous evaluations of emerging DNA sequence-based deep learning models that account for long-range dependencies.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cheng, W., Song, Z., Zhang, Y., Wang, S., Wang, D., Yang, M., Li, L., Ma, J.. 2025-01-08. DNALONGBENCH: A Benchmark Suite for Long-Range DNA Prediction Tasks. https://doi.org/10.1101/2025.01.06.631595
Cite the original work for its findings. Save a collection to share your selection of sources.