bioRxiv Science⌕ Search

Biology subjects

Ghent, S.

Publications and source records attributed to Ghent, S..

2 recordsLinked to original sources

Do AI Structure Predictors Capture Bound-State Disorder? A Benchmark on Fuzzy Protein Complexes

Fuzzy protein complexes, in which an intrinsically disordered protein (IDP) retains conformational disorder upon binding, pose a fundamental challenge for structure predictors trained on ordered systems, where crystal structures capture only the most ordered ensemble snapshot, making standard benchmarking metrics misleading. Here, we present the first systematic evaluation of AlphaFold3 (AF3), AlphaFold2-Multimer (AF2MM), Chai-1, and Boltz-2 on a curated dataset of fuzzy complexes from FuzDB, benchmarked against DockQ against PDB structures and NOE violation rates against manually curated BMRB restraint files, the first comprehensive collection of this kind. Across all four predictors, approximately 30% of NOE restraints were violated with nearly identical distributions regardless of predictor architecture or training data. DockQ scores fell uniformly within the Acceptable range, with AF3 marginally higher but exhibiting NOE violation rates equivalent to the weakest-performing model. Ensemble-level analysis using a first-principles implementation of the Hadzi thermodynamic model revealed that AF3 uniquely achieves near-zero mean helicity bias, in contrast to systematic overconfidence in the other predictors, yet all four models show poor per-residue helicity correlation with thermodynamic expectations. DockQ rankings reflect training data similarity to crystal structures rather than physical accuracy, and no current predictor captures fuzzy complex ensemble behavior. The FuzzyBench-NOE dataset, comprising NOE restraint files, predicted structures, interface hotspot annotations, and Hadzi-DSSP analysis outputs, is released on Zenodo (https://doi.org/10.5281/zenodo.20470556). Significance StatementNo benchmark exists for fuzzy protein complexes, where IDPs retain disorder upon binding. We show that four state-of-the-art structure predictors violate 30% of experimental NMR distance restraints invariantly regardless of architecture, while DockQ, the standard metric, is entirely uncorrelated with this failure. Ensemble-level analysis using the Ha[d]zi thermodynamic model reveals systematic helicity overconfidence across all predictors. Taken together, our findings imply that standard geometric metrics are fundamentally misleading for disordered systems, thus necessitating ensemble-aware evaluation.

biophysics↗

AlphaFold3 and Intrinsically Disordered Proteins: Reliable Monomer Prediction, Unpredictable Multimer Performance

AlphaFold3 represents a major advance in protein structure prediction, yet its performance on intrinsically disordered proteins remains uncharacterized. We present the first systematic evaluation of AF3 on disordered systems, revealing a striking dichotomy. For monomers, AF3s pLDDT scores reliably predict disorder (MCC: 0.693), matching AlphaFold2 and rivaling dedicated predictors. This consistency across fundamentally different architectures confirms that disorder prediction emerges from training data, not model design. For multimers, the picture grows complex. Despite comparable aggregate performance (mean DockQ: 0.563 vs 0.571), AF3 and AF2 achieve these results through fundamentally different mechanisms. Conventional structural features explain 58% of AF2s variance but only 42% of AF3s. Users cannot predict when AF3 will succeed or fail from interface properties alone. On disorder-to-order transitions (MFIB benchmark), both models perform equally well, successfully predicting final folded states. Yet seed variance analysis reveals AF3s failures are deterministic: the model converges to identical structures across independent runs, whether correct or incorrect, indicating rigid structural priors override available information. Our findings establish AF3 as reliable for the prediction of monomer disorder but unpredictable for multimers. Architectural innovation alone cannot overcome training data bias. Progress demands disorder-enriched datasets and ensemble sampling, not merely novel architectures.

bioinformatics↗