bioRxiv Science⌕ Search

Biology subjects

Klivans, A.

Publications and source records attributed to Klivans, A..

3 recordsLinked to original sources

Folding scFv--Antigen Complexes at Scale

AO_SCPLOWBSTRACTC_SCPLOWAccurate modeling of antibody-antigen (Ab-Ag) complexes is central to biologic development, yet the reliability and failures of modern Ab-Ag folding pipelines remain poorly characterized. Single-chain variable fragments (scFvs) are thera-peutically important antibodies, but large-scale evaluations of structure prediction models on scFv-Ag complexes are largely lacking. We introduce a scalable bench-marking pipeline that generates large ensembles of scFv-Ag structure predictions by cofolding a curated subset of 3,800 Ab-Ag complexes from SAbDab using multiple state-of-the-art models under diverse inference-time settings. The resulting dataset, SCALE (scFv-Ag CompLex Ensembles), includes standardized scFv-Ag sequences and around 200,000 predicted complexes spanning different models, sampling strategies, and auxiliary inputs. Using SCALE, we evaluate model performance in recovering correct scFv-Ag interfaces and assess the ability of existing confidence metrics to select the best structure from prediction ensembles. We find that while confidence scores effectively distinguish easy from hard scFv-Ag complexes, they often fail to identify the highest-quality interface for a given target. Further analysis shows that near-correct interfaces typically appear in ensembles but at low frequency, and inference-time choices such as sampling, recycling, and using evolutionary or structural information are crucial for accurate scFv-Ag complex predictions. Dataset and analysis code are available at https://huggingface.co/datasets/ravishah1/SCALE

bioinformatics↗

Enzyme Classification via Semi-Supervised Functional ResidueLearning

Predicting enzymatic function from a protein sequence is a fundamental task in protein discovery and engineering. In this paper, we present Semi-supervised Learning for Enzyme Classification (SLEEC): a semi-supervised learning framework that learns a function-aware protein representation for Enzyme Commision (EC) number prediction. SLEEC achieves SOTA performance on standard bench-marks and provides interpretable, residue-level annotations. We further demonstrate that our framework is robust to benign sequence modifications routinely observed in protein engineering workflows- such as appending functional tags- a desirable property that current ML frameworks lack. Our main technical contribution is a multiple sequence alignment (MSA)-based data augmentation technique for discovering sparse residue activations within a given enzyme sequence.

bioengineering↗

Generating functional and multistate proteins with a multimodal diffusion transformer

Generating proteins with the full diversity and complexity of functions found in nature is a grand challenge in protein design. Here, we present ProDiT, a multimodal diffusion model that unifies sequence and structure modeling paradigms to enable the design of functional proteins at scale. Trained on sequences, 3D structures, and annotations for 214M proteins across the evolutionary landscape, ProDiT generates diverse, novel proteins that preserve known active and binding site motifs and can be successfully conditioned on a wide range of molecular functions, spanning 465 Gene Ontology terms. We introduce a diffusion sampling protocol to design proteins with multiple functional states, and demonstrate this protocol by scaffolding enzymatic active sites from carbonic anhydrase and lysozyme to be allosterically deactivated by a calcium effector. Our results showcase ProDiTs unique capacity to satisfy design specifications inaccessible to existing generative models, thereby expanding the protein design toolkit.

bioinformatics↗