bioRxiv Science⌕ Search

Biology subjects

Frey, N. C.

Publications and source records attributed to Frey, N. C..

4 recordsLinked to original sources

Lab-in-the-loop therapeutic antibody design with deep learning

Therapeutic antibody design is a complex multi-property optimization problem with substantial promise for improvement with the application of machine-learning methods. Towards realizing that promise, we introduce "Lab-in-the-loop," a new approach that orchestrates state-of-the-art repertoire mining methods, generative machine learning models, multi-task property predictors, active learning ranking and selection, and in vitro experimentation in a semi-autonomous, iterative optimization loop. By automating the design of antibody variants, property prediction, ranking and selection of designs to assay in the lab, and ingestion of in vitro data, we enable an end-to-end approach to developing computationally-informed therapeutic antibody design pipelines. We apply lab-in-the-loop to eleven seed antibodies obtained via animal immunization with four clinically relevant antigen targets: EGFR, IL-6, HER2, and OSM. Over 1,800 unique antibody variants are tested throughout four rounds of iterative optimization identifying 3-100x better binding variants for all targets and 10/11 seeds, with the best binders exceeding 100 pM affinity, demonstrating a process by which end-to-end machine learning can be developed for therapeutic antibody development.

bioengineering↗

DyAb: sequence-based antibody design and property prediction in a low-data regime

Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a new model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in protein properties, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across molecules targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates novel antibody candidates with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for engineering therapeutic protein properties in low data regimes common in early-stage drug development.

bioengineering↗

Cramming Protein Language Model Training in 24 GPU Hours

Protein language models (pLMs) are ubiquitous across biological machine learning research, but state-of-the-art models like ESM2 take hundreds of thousands of GPU hours to pre-train on the vast protein universe. Resource requirements for scaling up pLMs prevent fundamental investigations into how optimal modeling choices might differ from those used in natural language. Here, we define a "cramming" challenge for pLMs and train performant models in 24 hours on a single GPU. By re-examining many aspects of pLM training, we are able to train a 67 million parameter model in a single day that achieves comparable performance on downstream protein fitness landscape inference tasks to ESM-3B, a model trained for over 15, 000x more GPU hours than ours. We open source our library1 for training and inference, LBSTER: Language models for Biological Sequence Transformation and Evolutionary Representation.

bioengineering↗

EquiFold: Protein Structure Prediction with a Novel Coarse-Grained Structure Representation

Designing proteins to achieve specific functions often requires in silico modeling of their properties at high throughput scale and can significantly benefit from fast and accurate protein structure prediction. We introduce EquiFold, a new end-to-end differentiable, SE(3)-equivariant, all-atom protein structure prediction model. EquiFold uses a novel coarse-grained representation of protein structures that does not require multiple sequence alignments or protein language model embeddings, inputs that are commonly used in other state-of-the-art structure prediction models. Our method relies on geometrical structure representation and is substantially smaller than prior state-of-the-art models. In preliminary studies, EquiFold achieved comparable accuracy to AlphaFold but was orders of magnitude faster. The combination of high speed and accuracy make EquiFold suitable for a number of downstream tasks, including protein property prediction and design.

bioinformatics↗