bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.08.12.744534

Active Learning Enables Efficient Directed Evolution of a Far-Red Fluorescent Protein with Minimal Experimental Data

Abstract

Fluorescent proteins are fundamental tools for cellular imaging. Most fluorescent proteins in routine use, including GFP, are derived from the jellyfish Aequorea victoria and emit blue-green light, which is strongly absorbed and scattered by tissue, limiting imaging depth. Far-red and near-infrared fluorescent proteins, engineered from bacteriophytochromes, address this limitation because far-red light penetrates tissue considerably further. However, these proteins are typically much dimmer than their A. victoria -derived counterparts. Improving brightness by conventional directed evolution requires screening large random mutant libraries, a process that is slow, labor-intensive, and often impractical outside specialized laboratories. We utilized an active-learning-guided directed evolution workflow that identified improved variants from substantially less data than conventional screening. Each round coupled automated, miniaturized cell-free protein expression directly from a DNA template without cloning or cell culture, with a machine-learning model retrained on cumulative sequence-function data to nominate the most informative variants for the next round. Applied to miRFP670nano3, this workflow screened 120 variants across successive rounds and identified twelve with improved brightness, the best four-fold brighter in bacterial systems. However, these gains did not translate when the variants were evaluated in mammalian cells, indicating that performance can be strongly dependent on cellular context. Retrospective simulation across benchmark datasets from ProteinGym showed that performing more experimental batches with fewer samples per batch consistently accelerated convergence to high-fitness sequences. Incorporating protein-language-model derived zero-shot fitness priors also accelerated convergence, but only in proportion to how well each prior score correlated with the true fitness landscape. Together, these findings established generalizable design rules, favoring smaller acquisition batches and confidence-weighted priors, for engineering proteins from minimal experimental data. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=55 SRC="FIGDIR/small/744534v1_ufig1.gif" ALT="Figure 1"> View larger version (11K): org.highwire.dtl.DTLVardef@14cd238org.highwire.dtl.DTLVardef@7d8334org.highwire.dtl.DTLVardef@30e9f2org.highwire.dtl.DTLVardef@14f3a0c_HPS_FORMAT_FIGEXP M_FIG C_FIG

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Brown, D. V., Cross, R. S., Zhu, S., Hill, T., Sok, C. L., Jenkins, M. R., Dramicanin, M., Bowden, R.. 2026-08-17. Active Learning Enables Efficient Directed Evolution of a Far-Red Fluorescent Protein with Minimal Experimental Data. https://doi.org/10.64898/2026.08.12.744534

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Gene expression noise is reduced in communicating synthetic cell populations

A major goal in bottom-up synthetic biology is the construction of multicellular synthetic systems capable of coordinated and robust collective behaviours. However, robustness is often limited by noise and variability arising from increased molecular complexity. Whilst communication has been implemented in synthetic multi-cellular systems, the ability for communication to suppress cell free gene expression variability in populations of synthetic cells remain unexplored. To address this, we encapsulated the Lux and Las quorum sensing gene circuits in lipid vesicles under cell-free conditions to test the effect of communication on reducing cell-free gene expression variability across the population. Our results show that communication, limiting expression resources, and membrane surface effects can reduce gene expression variability. Resource limited Gillespie simulations for transcription and translation show that communication-mediated coupling reduces population-level expression noise under constrained and excess resource conditions. Together, our work provides simple strategies to reduce gene expression variability and thereby improve robustness in synthetic multicellular systems, an important criteria for the future applications of synthetic cells.

synthetic biology↗

Boolean Logic-responsive FRET Biosensors via Genetically Encoded Autonomous Compilation

Forster resonance energy transfer (FRET) is commonly used to monitor protein-protein interactions in situ. The high spatiotemporal resolution and facile implementation inside complex molecular environments have spearheaded FRET's widespread adoption in biosensing. Despite these advantages, current FRET biosensors are largely restricted to the detection of the presence/absence of individual inputs and are thus unable to sense several multiplexable inputs simultaneously within complex milieu of biological environments. In this work, we introduce a generalizable strategy to construct genetically encoded protein-based FRET biosensors capable of recognizing multiple inputs following Boolean logic-type (YES/OR/AND) operations. These topologically specified FRET sensors powerfully expand the input capacity in sensing protein-protein interactions while providing a user-programmable platform for monitoring heterogeneous biological activities both in vitro and in living cells.

synthetic biology↗

AI-Guided Multi-Objective Engineering of Glucoamylase Enables Acidification-Free Starch Saccharification

Glucoamylase is essential for industrial starch saccharification, but the limited thermostability and near-neutral pH tolerance of fungal glucoamylases necessitate cooling and acidification of liquefied starch. Here, we developed an artificial intelligence-guided strategy to simultaneously improve the thermostability, pH tolerance, and catalytic activity of glucoamylase from Penicillium oxalicum (PoGA). Two property-specific machine-learning models, CASPE-T and CASPE-A, identified substitutions associated with thermostability and pH tolerance, respectively. Experimental screening identified beneficial substitutions in 11 of 21 CASPE-T and 12 of 22 CASPE-A candidates. Folding-energy-guided recombination integrated the two traits while maintaining structural compatibility. The optimal variant, PoGA T513E/Q305N, exhibited 2.21-fold higher specific activity than the wild type, with half-life extended from 22.3 to 57.9 min at 60 degrees C and from 16.6 to 64.7 min at pH 8.0. Molecular dynamics simulations attributed these improvements to reinforcement of high-occupancy hydrogen-bonding networks, suppression of conformational fluctuations in the linker and carbohydrate-binding module, enhanced long-range dynamic coordination, and preservation of a compact catalytic architecture. At 60 degrees C and pH 6.5 without acidification, PoGA T513E/Q305N produced 219.9 g/L glucose and achieved 89.1% starch conversion, 31.4% higher than the wild type. This work provides an efficient framework for multi-objective enzyme engineering and sustainable starch biorefining.

synthetic biology↗