Machine learning-guided optimization of p-coumaric acid production in yeast
Industrial biotechnology uses Design-Build-Test-Learn (DBTL) cycles to accelerate the development of microbial cell factories, required for the transition to a bio-based economy. To use them effectively, appropriate connections between each phase of the cycle are crucial. Using p-coumaric acid production in Saccharomyces cerevisiea as case study, we propose the use of one-pot library generation, random screening, targeted sequencing and machine learning (ML) as links during DBTL cycles. We showed that the robustness and flexibility of ML models strongly enable pathway optimization, and propose feature importance and SHAP values as a guide to expand the design space of original libraries. This approach led to a 68% increased production of p-coumaric acid within two DBTL cycles.