bioRxiv Science⌕ Search

Biology subjects

Steinberg, D. M.

Publications and source records attributed to Steinberg, D. M..

2 recordsLinked to original sources

Data-Efficient Exploration of Enzyme Function Using Family-Specific Machine Learning

Enzymes are essential biocatalysts across diverse industries, driving demand for high-performing variants. Foundation models are attractive for guiding enzyme discovery, but often lack the resolution to model subtle variations driving function within homologous families. Navigating these rugged functional landscapes to identify elite variants remains challenging and experimentally costly, even when guided by such models. Here we show that coupling dense, family-specific experimental screening with targeted, sequence-based deep learning provides a data-efficient discovery strategy. We experimentally screened 1,513 natural homologues from an esterase superfamily (>7,500 assays) and used this functional landscape to train task-specific models that predict activity, thermostability, and substrate specificity from sequence alone. Prospective experimental validation of previously untested sequences demonstrated that these task-specific models significantly outperformed generalist pre-trained and physics-based models in enriching for target traits. Residue-level attribution further indicated that the models captured sequence patterns consistent with underlying structural features. Finally, retrospective simulations showed that iterative retraining compresses the search space, discovering 60% of top-tier hits using nearly half the samples required by pre-trained baseline models. Together, these results highlight that machine learning can provide mechanistic insight, and that integrating targeted data acquisition with iterative machine learning provides a more data-efficient discovery strategy than relying on generic model scale.

bioengineering↗

Ultrahigh throughput screening to train generative protein models for engineering specificity into unspecific peroxygenases

Enzyme engineering plays a vital role in tailoring biocatalyst performance to meet the needs of target applications. However, the number of sequence trajectories possible from a single wildtype enzyme sequence is too vast to traverse experimentally. Here we present a novel approach that first expands the experimentally accessible sequence space using ultrahigh throughput screening (uHTS), and then uses indirect and low fidelity assay data to create a "fingerprint" for a target enzyme class. Experimental data from microfluidic uHTS are extracted and used to engineer specificity into an unspecific peroxygenase (UPO) from Aspergillus brasiliensis (AbrUPO). We created a library with more than 5 million different variants expressed in Komagataella phaffii (Pichia pastoris). Microfluidic droplet sorting was then used to generate a dataset of >30,000 unique sequences paired with function data. This dataset was then used to train a task-specific generative model using the Variational Search Distributions (VSD) framework. We compared the variants selected by rank aggregation from the screening data (R series) with novel sequences generated by the refined generative model (G series). While the wildtype enzyme produces nearly equal amounts of both the desired styrene oxide and undesired phenylacetaldehyde products, three out of five of the highest scoring G series variants produced product mixtures more enriched in the desired compound. In comparison, only one of the five highest scoring R series variants showed this improvement. Overall, the variant most enriched in desired product, G929, produced 2.4x more styrene oxide than phenylacetaldehyde, while G3 and G167, produced the highest quantities of desired product at 2.3x enrichment over the undesired product. Further analysis confirmed that our task-specific generative model outperforms existing models pre-trained on large publicly available datasets. This unique combination of uHTS and generative protein modelling provides an intelligent exploration mechanism which not only enables efficient enzyme discovery, but also accelerates optimization and enables predictive insights that are difficult to achieve with either approach alone.

bioengineering↗