bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.07.02.736085

A Comprehensive Evaluation of Protein Structure Prediction Models for Short Peptides

Abstract

Short peptides pose distinct challenges for computational structural biology due to their lack of stable tertiary structures, high conformational flexibility, and limited evolutionary signals. To address how modern deep-learning architectures navigate these challenges, we conducted a comprehensive benchmarking of five state-of-the-art protein structure prediction models: AlphaFold2, RoseTTAFold2, ESMFold, OmegaFold, and DMPfold2. Using a curated dataset of experimentally determined short peptide structures (10-49 amino acids) from the Protein Data Bank, we systematically evaluated predictive performance across varying sequence lengths and secondary structure classes. Our results demonstrate that prediction accuracy systematically improves with peptide length. Furthermore, all models perform significantly better on -helical and mixed-structure peptides compared to {beta}-sheet-rich and intrinsically disordered sequences. Among the evaluated methods, AlphaFold2 and the single-sequence language models, ESMFold and Omegafold proved to be the most consistent and accurate overall. We also observed that internal model confidence scores are imperfectly calibrated for short peptides, necessitating cautious interpretation. Finally, by extending our analysis to the dbAMP3 dataset of uncharacterized antimicrobial peptides, we demonstrate that a multi-model consensus approach provides a rational framework for identifying robust structural hypotheses in the absence of experimental reference structures.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ghosh, B., MUKHERJEE, A.. 2026-07-03. A Comprehensive Evaluation of Protein Structure Prediction Models for Short Peptides. https://doi.org/10.64898/2026.07.02.736085

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Scaling of structural variability of ecDNA polymer condensates with copy number boosts and stabilises oncogene regulatory contacts

Extrachromosomal DNAs (ecDNAs) form highly heterogeneous condensates in cancer cells that drive oncogene overexpression, yet how structural variability coexists with stable gene regulation remains unclear. Here, we develop a minimal polymer physics model of MYC-harbouring COLO320-DM ecDNAs, where BRD4-like complexes bind and bridge cognate sites along ecDNA rings. Above a critical binder concentration, ecDNAs phase separate into condensates exhibiting diverse conformations because of their thermodynamic folding degeneracy. Despite this variability, condensates retain conserved interaction scaffolds that give rise to reproducible contact patterns, including in-trans associated domains (I-TADs), genomic regions enriched in intermolecular regulatory contacts between distinct ecDNAs. We find that condensate 3D architecture follows universal scaling relations with ecDNA copy number, n, remaining robust to model parameter changes. Regulatory contacts within I TADs increase linearly with n, yet they are one order of magnitude stronger than in size matched control regions outside I TADs, whereas their relative fluctuations are markedly suppressed as n increases. This scaling produces enhanced, low-noise regulatory environments for oncogenes embedded within I-TADs, such as PVT1-MYC fusions, whereas the canonical MYC copy, located outside, is less amplified as experimentally observed. Our findings reveal universal polymer physics principles underlying ecDNA condensate organization, offering a mechanistic basis for selective oncogene amplification and potential advantages in cancer progression.

biophysics↗

High-resolution mapping of RNA structural maturation during Cas9 assembly with ABEL-FRET

The structural flexibility of RNA is essential for forming ribonucleoprotein (RNP) complexes, which regulate diverse biological processes. This intrinsic property permits RNA to act as a dynamic scaffold along the assembly pathway as it folds into a specific structure for initial recognition by protein and undergoes conformational rearrangements for functional maturation as a complex. Yet, RNA flexibility and RNP multicomponent assembly create significant obstacles for traditional structural methods. To overcome these challenges, we applied recently developed ABEL-FRET spectroscopy to measure tether-free single-molecule Forster resonance energy transfer (smFRET) over extended observation times. Furthermore, ABEL-FRET enables the unique ability for simultaneous measurements of ultrahigh resolution smFRET and hydrodynamic size of individual complexes, which offers distinct advantages for studying dynamic RNA molecules that undergo assembly via sequential binding events. Using ABEL-FRET, we explored how the guide RNA (gRNA) of CRISPR genome editing system folds and modulates its structural flexibility to carry out the roles required for each assembly state from its unbound apo form to the functional Cas9 RNP state for target DNA cleavage. Multi-perspective view of gRNA structure gained by probing its two primary functional domains enabled to capture dramatic changes in gRNA flexibility that are highly dependent on its specific structural domains as well as assembly states. Collectively, our work with ABEL-FRET highlights the intrinsic link between the structural flexibility of RNA and its functionality in RNP assembly.

biophysics↗

De novo design of functional RNAs through higher-order interactions

Designing RNA sequences that reliably adopt functional three-dimensional structures remains a central challenge in RNA engineering because folding depends on cooperative interactions beyond canonical base pairing. Here we present DS3dRNA, an interaction-based framework for de novo RNA sequence design that combines a three-body statistical potential with physics-guided sequence sampling and supports design against multiple conformations. Across the evaluated benchmarks, DS3dRNA outperformed representative RNA inverse-design methods in native-sequence recovery and agreement between predicted and target structures. Energy-sequence-quality analyses further showed that lower design energies generally accompanied higher sequence recovery and macro-averaged F1 scores (MacroF1). Experimentally tested Mango II designs retained high-affinity fluorogenic activity, and five twister ribozyme designs yielded mean endpoint cleavage fractions of 37.7-50.6%, compared with 23.5% for the wild type. These results establish explicit higher-order interaction scoring as a complementary approach to emerging data-driven RNA design methods and provide a framework for designing functional RNAs from experimental or predicted structural ensembles.

biophysics↗