bioRxiv Science⌕ Search

Biology subjects

Fahlberg, S. A.

Publications and source records attributed to Fahlberg, S. A..

3 recordsLinked to original sources

Functional alignment of protein language models via reinforcement learning

Protein language models (pLMs) enable generative design of novel protein sequences but remain fundamentally misaligned with protein engineering goals, as they lack explicit understanding of function and often fail to improve properties beyond those found in nature. We introduce Reinforcement Learning from eXperimental Feedback (RLXF), a general framework that aligns protein language models with experimentally measured functional objectives, drawing inspiration from the methods used to align large language models like ChatGPT. Applied across five diverse protein families, RLXF improves generation of high-functioning variants beyond pre-trained baselines. We demonstrate this with CreiLOV, an oxygen-independent fluorescent protein, where RLXF-aligned models generate sequences with significantly enhanced fluorescence, including the most fluorescent CreiLOV variants reported to date. Our results indicate that RLXF-aligned models effectively integrate the evolutionary knowledge encoded in pre-trained pLMs with experimental observations, improving the success rate of generated sequences and enabling the discovery of synergistic mutation combinations that are difficult to identify through zero-shot or evolutionary approaches. RLXF provides a scalable and accessible approach to steer generative models toward desired biochemical properties, enabling function-driven protein design beyond the limits of natural evolution.

bioengineering↗

Neural network extrapolation to distant regions of the protein fitness landscape

Machine learning (ML) has transformed protein engineering by constructing models of the underlying sequence-function landscape to accelerate the discovery of new biomolecules. ML-guided protein design requires models, trained on local sequence-function information, to accurately predict distant fitness peaks. In this work, we evaluate neural networks capacity to extrapolate beyond their training data. We perform model-guided design using a panel of neural network architectures trained on protein G (GB1)-Immunoglobulin G (IgG) binding data and experimentally test thousands of GB1 designs to systematically evaluate the models extrapolation. We find each model architecture infers markedly different landscapes from the same data, which give rise to unique design preferences. We find simpler models excel in local extrapolation to design high fitness proteins, while more sophisticated convolutional models can venture deep into sequence space to design proteins that fold but are no longer functional. Our findings highlight how each architectures inductive biases prime them to learn different aspects of the protein fitness landscape.

synthetic biology↗

Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production

Fatty acyl reductases (FARs) catalyze the reduction of thioesters to alcohols and are key enzymes for the microbial production of fatty alcohols. Many existing metabolic engineering strategies utilize these reductases to produce fatty alcohols from intracellular acyl-CoA pools; however, acting on acyl-ACPs from fatty acid biosynthesis has a lower energetic cost and could enable more efficient production of fatty alcohols. Here we engineer FARs to preferentially act on acyl-ACP substrates and produce fatty alcohols directly from the fatty acid biosynthesis pathway. We implemented a machine learning-driven approach to iteratively search the protein fitness landscape for enzymes that produce high titers of fatty alcohols in vivo. After ten design-test-learn rounds, our approach converged on engineered enzymes that produce over twofold more fatty alcohols than the starting natural sequences. We further characterized the top identified sequence and found its improved alcohol production was a result of an enhanced catalytic rate on acyl-ACP substrates, rather than enzyme expression or KM effects. Finally, we analyzed the sequence-function data generated during the enzyme engineering to identify sequence and structure features that influence fatty alcohol production. We found an enzymes net charge near the substrate-binding site was strongly correlated with in vivo activity on acyl-ACP substrates. These findings suggest future rational design strategies to engineer highly active enzymes for fatty alcohol production.

synthetic biology↗