ML4SD: Leveraging Machine Learning and High-Throughput Search Algorithms for an Iterative Growth-Coupled Design Innovation
Optimizing microbial biomanufacturing is required if renewable and waste carbon are to replace petrochemical routes at competitive titers, rates, and yields. Growth-coupled (GC) production supports that goal by linking target synthesis to biomass formation, so product formation is required for growth. Constructing knockout strains yielding GC production from a list of candidate genes is labor and time demanding. This results in few in vivo tested designs, which hampers standard machine-learning methods to learn GC patterns for specific bioprocesses. We therefore developed ML4SD, an active-learning Design-Build-Test-Learn (DBTL) cycle that trains ensembles on genome-scale metabolic model (GEM) scores of knockout designs, sampling the next designs from predicted model performance and error. That cycle generalizes only if the initial library is large and diverse, including suboptimal and non-viable designs; libraries restricted to minimal designs or Pareto-optimal knockouts were found to generate models overfitting. To meet those specific demands a novel strain design algorithm, gcSwarms, was developed and tested for a diverse set of bioprocesses within Pseudomonas putida iJN1462. ML4SD was tested with an in silico case study converting lignin-derived 4-hydroxybenzoate to 6-caprolactam, the nylon-6 monomer. ML4SD results showed improvements of up to 164% on carbon yield, recovering a shared SHAP motif that redirects TCA flux through acetyl-CoA. Importantly it reaches that result using 2.5- to 7.1-fold fewer designs than a gcSwarms-only search, demonstrating the data efficiency of this method.