bioRxiv Science⌕ Search

Biology subjects

Plaikner, A.

Publications and source records attributed to Plaikner, A..

2 recordsLinked to original sources

Protein Language Models: Learning From Evolution, Designing Beyond It

Protein language models (PLMs) have transformed our ability to learn from evolutionary sequence space, but protein engineering ultimately asks a different question: not what evolution selected, but what we should build next. Zero-shot likelihoods therefore provide useful, but not universal, measures of fitness and can misalign with engineering objectives. Experimental supervision redirects these priors towards properties of interest, enabling target-specific prediction and closed-loop optimization. PLMs thereby complement structure-based design: structural methods provide geometric control, while PLMs integrate experimental feedback to optimize functional and developability properties. Realizing this potential requires evaluation beyond retrospective predictive accuracy, towards extrapolation, multi-objective optimization, and prospective experimental success. We argue for recurring competitions combining scalable predictive benchmarks with prospective experimental challenges. Ultimately, progress should be measured not by predicting existing experiments, but by enabling successful new ones.

bioinformatics↗

Pseudoperplexity Probes Memorization in Protein Language Models

Protein Language Models (pLMs) have significantly advanced computational biology. Yet their scale and reliance on redundant training data raise a fundamental question: do pLMs generalize the statistical grammar of proteins, or do they simply memorize their training data? To investigate this, we used pseudoperplexity as a probe for sequence-level memorization, comparing ProtT5's pseudoperplexity on a pre-training proxy dataset against a post-training holdout of genuinely novel sequences. To ensure a valid comparison, we matched the datasets by sequence length, cluster size, and taxonomic family. As a statistical baseline, we trained n-gram language models; analysis of higher-order n-gram composition and a statistically significant divergence in perplexity confirmed that the post-training sequences were genuinely novel at the local sequence level. ProtT5 showed a statistically significant difference in pseudoperplexity between seen and unseen sequences, though further analysis revealed this memorization signal to be modest. These findings suggest that ProtT5 exhibits detectable but limited memorization of its training data as measured by a pseudoperplexity-based probe.

bioinformatics↗