bioRxiv Science⌕ Search

Biology subjects

Gomez-Uribe, C.

Publications and source records attributed to Gomez-Uribe, C..

1 recordsLinked to original sources

Designing diverse and high-performance proteins with a large language model in the loop

We present a novel protein engineering approach to directed evolution with machine learning that integrates a new semi-supervised neural network fitness prediction model, Seq2Fitness, and an innovative optimization algorithm, biphasic annealing for diverse adaptive sequence sampling (BADASS) to design sequences. Seq2Fitness leverages protein language models to predict fitness landscapes, combining evolutionary data with experimental labels, while BADASS efficiently explores these landscapes by dynamically adjusting temperature and mutation energies to prevent premature convergence and find diverse high-fitness sequences. Seq2Fitness predictions improve the Spearman correlation with fitness measurements over alternative model predictions, e.g., from 0.34 to 0.55 for sequences with mutations residues that are absent from the training set. BADASS requires less memory and computation compared to gradient-based Markov Chain Monte Carlo methods, while finding more higher-fitness sequences and maintaining sequence diversity in protein design tasks for two different protein families with hundreds of amino acids. For example, for both protein families 100% of the top 10,000 sequences found by BADASS have higher Seq2Fitness predictions than the wildtype sequence, versus a broad range between 3% to 99% for competing approaches with often many fewer than 10,000 sequences found. The fitness predictions for the top, top 100th, and top 1,000th sequences found by BADASS are all also higher. In addition, we developed a theoretical framework to explain where BADASS comes from, why it works, and how it behaves. Although we only evaluate BADASS here on amino acid sequences, it may be more broadly useful for exploration of other sequence spaces, including DNA and RNA. To ensure reproducibility and facilitate adoption, our code is publicly available here. Author summaryDesigning proteins with enhanced properties is essential for many applications, from industrial enzymes to therapeutic molecules. However, traditional protein engineering methods often fail to explore the vast sequence space effectively, partly due to the rarity of high-fitness sequences. In this work, we introduce BADASS, an optimization algorithm that samples sequences from a probability distribution with mutation energies and a temperature parameter that are updated dynamically, alternating between cooling and heating phases, to discover high-fitness proteins while maintaining sequence diversity. This stands in contrast to traditional approaches like simulated annealing, which often converge on fewer and lower fitness solutions, and gradient-based Markov Chain Monte Carlo (MCMC), also converging on lower fitness solutions and at a significantly higher computational and memory cost. Our approach requires only forward model evaluations and no gradient computations, enabling the rapid design of high-performing proteins that can be validated in the lab, especially when combined with our Seq2Fitness models. BADASS represents a significant advance in computational protein engineering, opening new possibilities for diverse applications.

bioengineering↗