bioRxiv Science⌕ Search

Biology subjects

Lahman, M. C.

Publications and source records attributed to Lahman, M. C..

2 recordsLinked to original sources

The Synthetic Epitope Atlas: High-Throughput Design and Validation of De Novo Antibody-Antigen Complexes

AO_SCPLOWBSTRACTC_SCPLOWDe novo antibody design models lack sufficient training data to reliably generalize. We demonstrate scalable generation of structural training data for machine learning-driven antibody design by linking in silico designs of antibody-antigen complexes to high-throughput experimental binding validation. Using AlphaSeq, a yeast-based platform for measuring protein binding affinities, we measure the affinity and specificity of thousands of de novo "synthetic epitope proteins" (SEPs) designed to bind to VHHs. The resulting Synthetic Epitope Atlas (SEPIA) pairs over 26 million on- and off-target affinity measurements with computationally designed VHH-SEP "pseudo-structures." We validate strong, specific binding for 1,161 pseudo-structures and >75,000 VHH and SEP mutational variants. We show that these pseudo-structures complement existing structural databases and enable ML models to outperform confidence metrics commonly used to rank de novo antibody designs. Taken together, SEPIA establishes a scalable framework for improving de novo antibody design by augmenting sparse structural data with large-scale experimental binding data.

synthetic biology↗

AlphaBind, a Domain-Specific Model to Predict and Optimize Antibody-Antigen Binding Affinity

Antibodies are versatile therapeutic molecules that utilize combinatorial sequence diversity to cover a vast fitness landscape. However, designing optimal antibody sequences remains a major challenge. Recent advances in deep learning provide opportunities to address this challenge by learning sequence-function relationships to accurately predict fitness landscapes. These models enable efficient in silico prescreening and optimization of antibody candidates. By focusing experimental efforts on the most promising candidates guided by deep learning predictions, antibodies with optimal properties can be designed more quickly and effectively. Here we present AlphaBind, a domain-specific model that utilizes protein language model embeddings and pre-training on millions of quantitative laboratory measurements of antibody-antigen binding strength to achieve state-of-the-art performance for guided affinity optimization of parental antibodies. We demonstrate that an AlphaBind-powered antibody optimization pipeline can deliver candidates with substantially improved binding affinity across four parental antibodies (some of which were already affinity-matured) and using two different types of training data. Resulting candidates, ranging up to 11 mutations from parental sequence, yield a sequence diversity that allows for optimization of other biophysical characteristics, all while using only a single round of data generation for each parental antibody. AlphaBind weights and code are publicly available at: https://github.com/A-Alpha-Bio/alphabind.

synthetic biology↗