bioRxiv Science⌕ Search

Biology subjects

si, y.

Publications and source records attributed to si, y..

2 recordsLinked to original sources

NUMonomer enables accurate and scalable nucleic acid structure prediction from primary sequence alone

Accurate and efficient prediction of three-dimensional nucleic acid structures can accelerate functional characterization and enable downstream applications. Recent deep-learning methods have substantially improved nucleic acid structure prediction by incorporating auxiliary inputs such as multiple sequence alignments, secondary-structure annotations, and representations from pretrained language models. However, prediction accuracy remains limited, and generating these auxiliary inputs can be computationally expensive. Here we show that learning the hierarchical organization of experimentally determined structures across multiple scales, from recurring local conformations to global fold topologies, together with exploiting representations shared between RNA and single-stranded DNA, improves model generalization. Guided by these findings, we developed NUMonomer, an end-to-end deep-learning framework trained with input sequences spanning thousands of nucleotides on a joint RNA and single-stranded DNA dataset to predict nucleic acid structures directly from sequence. Despite requiring no auxiliary inputs, NUMonomer matches or outperforms leading prediction methods on benchmarks comprising CASP16 RNA targets and non-redundant sets of experimentally determined RNA and single-stranded DNA structures, with particularly pronounced improvements for longer RNAs. Its efficient and scalable architecture also reduces inference costs by approximately two orders of magnitude relative to the evaluated methods, enabling large-scale structure prediction. Together, these findings provide insight into generalization in biomolecular structure learning and establish NUMonomer as a practical framework for nucleic acid structure prediction.

molecular biology↗

End-to-end single-stranded DNA sequence design with all-atom structure reconstruction

Designing biological sequences that fold into predefined conformations is a central challenge in bioengineering. Although deep learning has enabled significant advances in protein and RNA sequence design, progress in single-stranded DNA (ssDNA) design has been constrained by the limited availability of structural data. To address this challenge, we introduce InvDNA, a deep learning-based method that designs ssDNA sequences directly from backbone atomic coordinates. This end-to-end formulation avoids the loss of structural information during backbone-to-feature conversion and further accommodates flexible backbone representations, dynamic sequence masking, and structural reconstruction objectives. These strategies bolster InvDNAs ability to generalize across diverse ssDNA structural contexts while enabling additional functionalities, including generating diverse sequences for a given backbone, reconstructing nucleotide conformations from backbone and preserving functional sites. In benchmarks using experimentally determined ssDNA structures, InvDNA demonstrates more than a twofold improvement in sequence recovery compared with existing ssDNA and RNA sequence design approaches. Further computational validation using AlphaFold3 shows that 44.4% of InvDNA-designed sequences successfully fold into their predefined conformations. Notably, this success rate increases when backbone coordinates are perturbed to diversify the InvDNA-designed sequences. Collectively, these results establish InvDNA as a robust framework for rational ssDNA engineering.

molecular biology↗