bioRxiv Science⌕ Search

Biology subjects

Blalock, N.

Publications and source records attributed to Blalock, N..

4 recordsLinked to original sources

Integrating complementary biological information for multi-objective enzyme engineering

Enzyme catalysts are increasingly used for sustainable pharmaceutical manufacturing, but engineering industrial biocatalysts remains challenging as multiple catalytic and developability properties must be optimized simultaneously from limited experimental data. Here, we develop a machine learning-guided multi-objective design framework that integrates sparse functional measurements with complementary evolutionary and structural information to engineer the ketoreductase Gre2. The resulting designs achieved simultaneous improvements in catalytic performance, protein yield, and thermal stability, with selected designs retaining improved performance under process-relevant conditions. More broadly, our results demonstrate that integrating complementary biological information enables efficient multi-objective enzyme engineering from sparse experimental data, providing a general strategy for accelerating industrial biocatalyst development.

bioengineering↗

Machine learning-guided olivetolic acid cyclase engineering enables tailored cannabinoid biosynthesis in yeast

Cannabinoids comprise a diverse class of bioactive natural products with important therapeutic potential, but efficient microbial production remains limited by pathway bottlenecks and challenges in engineering key biosynthetic enzymes. Here, we develop a machine learning-guided approach to engineer olivetolic acid cyclase (OAC), a critical control point in cannabinoid biosynthesis that governs both pathway flux and product selectivity. We first generated sequence-function data from 152 CsOAC variants spanning homolog screening, recombination, and mutagenesis libraries. Using these measurements, we trained multi-task models to predict pathway-level production of olivetolic acid (OA), divarinic acid (DVA), and competing byproducts, together with a variational autoencoder that captured evolutionary constraints across the broader enzyme family. Across three rounds of iterative design and testing, this approach identified CsOAC variants that substantially increased production and selectivity of both OA and DVA. When introduced into engineered Yarrowia lipolytica strains, these variants enabled production of tetrahydrocannabinolic acid (THCA) and the minor cannabinoid tetrahydrocannabivarinic acid (THCVA) at titers exceeding previous yeast systems. Analysis of top-performing variants revealed mutations influencing substrate selectivity and catalytic performance, providing insight into the determinants of CsOAC function. More broadly, this work demonstrates how machine learning-guided enzyme engineering can improve pathway performance and expand access to major and minor cannabinoids through microbial biosynthesis.

bioengineering↗

Explicit representation of germline and non-germline residues improves antibody language modeling

Antibodies originate from germline templates and are diversified by somatic hypermutation, producing sequences in which conserved germline residues scaffold structure while rare non-germline (NGL) substitutions refine antigen binding. Current antibody language models (ALMs) treat all residues equivalently and inherit a germline bias that systematically down-weights functionally critical NGL mutations as statistical noise. We introduce PRISM, a germline-aware ALM that explicitly represents germline and nongermline residues as distinct token types over a factorized 53-token vocabulary. PRISM achieves state-of-the-art pseudo-perplexity in hypervariable CDRs and is uniquely positively correlated with experimental binding affinity across three deep mutational scanning landscapes on which all compared ALMs anti-correlate. The dual-vocabulary further enables property-specific controllable generation previously unattainable with entangled ALMs. NGL-directed sampling improves physics-based binding scores while GL-directed sampling preserves stability and solubility. These results establish disentangled germline/non-germline representation as a substantive advance in antibody language modeling.

immunology↗

Functional alignment of protein language models via reinforcement learning

Protein language models (pLMs) enable generative design of novel protein sequences but remain fundamentally misaligned with protein engineering goals, as they lack explicit understanding of function and often fail to improve properties beyond those found in nature. We introduce Reinforcement Learning from eXperimental Feedback (RLXF), a general framework that aligns protein language models with experimentally measured functional objectives, drawing inspiration from the methods used to align large language models like ChatGPT. Applied across five diverse protein families, RLXF improves generation of high-functioning variants beyond pre-trained baselines. We demonstrate this with CreiLOV, an oxygen-independent fluorescent protein, where RLXF-aligned models generate sequences with significantly enhanced fluorescence, including the most fluorescent CreiLOV variants reported to date. Our results indicate that RLXF-aligned models effectively integrate the evolutionary knowledge encoded in pre-trained pLMs with experimental observations, improving the success rate of generated sequences and enabling the discovery of synergistic mutation combinations that are difficult to identify through zero-shot or evolutionary approaches. RLXF provides a scalable and accessible approach to steer generative models toward desired biochemical properties, enabling function-driven protein design beyond the limits of natural evolution.

bioengineering↗