bioRxiv · 10.64898/2026.08.28.747907
Towards Sparse Causal Features for Zero-shot Mutation Effect Prediction in a Protein Language Model
Abstract
Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patching to identify the latent features that causally mediate zero-shot mutation effect prediction in ESM-2 650M. We evaluate this framework over 67 mutations ranging from strongly deleterious to weakly deleterious in the DNAJA1 J-domain, where ESM-2 predictions agree strongly with deep mutational scanning measurements. We find that circuits selected by indirect effect recover the model's predictions more efficiently and provide more informative biological explanations than those selected by raw activation changes, showing that activation magnitude does not necessarily reflect causal importance. We find that related substitutions reuse substantial portions of their recovered circuits, ranging from 40% to 75%, and that the shared features often represent residues in three-dimensional contact with the mutation site. To our knowledge, our work provides the first causal, feature-level account of zero-shot mutation effect prediction in a pLM.
Explore related subjects
Keep this discovery
Mohanty, S., Phutela, M., Green, A. G.. 2026-09-03. Towards Sparse Causal Features for Zero-shot Mutation Effect Prediction in a Protein Language Model. https://doi.org/10.64898/2026.08.28.747907
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.