bioRxiv · 10.64898/2026.06.04.730121
Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations
Abstract
Deep learning models have emerged as promising tools in protein engineering. In particular, they can be used to predict mutation fitness without the need for task-specific training, a process known as zero-shot prediction. However, the respective strengths and limitations of zero-shot predictions remain poorly understood. Here, we argue that commonly used large-scale benchmarks obscure important failure modes relevant to practical protein engineering, including an inability to pinpoint highly fit mutations or variants driving new-to-nature functions. Beyond these practical failures, we identify a fundamental limitation of zero-shot prediction: a generic fitness score cannot simultaneously optimize for distinct, competing engineering targets, meaning it is inherently disconnected from the phenotype of interest. Moreover, in a systematic comparison of a wide range of available models, we demonstrate that most models show comparable zero-shot performance, irrespective of model architecture and/or input modality (sequence vs. structure). Ultimately, we find that zero-shot predictions serve only as coarse filters separating fit mutations from deleterious ones, failing to reliably identify the mutations that would be most valuable in protein engineering.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Woolley, P. R., Feller, A., Ellington, A. O., Wilke, C. O.. 2026-06-07. Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations. https://doi.org/10.64898/2026.06.04.730121
Cite the original work for its findings. Save a collection to share your selection of sources.