bioRxiv Science⌕ Search

Biology subjects

Klakow, D.

Publications and source records attributed to Klakow, D..

2 recordsLinked to original sources

Generative AI designs functional thiolation domains for reprogramming non-ribosomal peptide synthetases

Large language models and generative protein design promise to accelerate biotechnology, but it remains unclear whether they can engineer dynamic megasynth(et)ases whose activity depends on transient, context-specific domain interfaces. Non-ribosomal peptide synthetases (NRPSs) are an especially demanding target, yet a high-value one because they produce many clinically important natural products and offer a route to analogs that are often difficult or impractical to access by chemical synthesis. Here we integrate pretrained generative models (ESM3, ProteinMPNN and EvoDiff) with design-build-test-learn cycles and data-guided prioritization to generate 76 de novo thiolation (T) domains. We built and tested 578 recombinant NRPS variants in vivo spanning minimal, full-length and hybrid assembly lines. AI-designed T-domains supported product formation across architectures, enabled catalytically active hybrids at recombined junctions and increased yields by up to [~]3-fold relative to NRPSs carrying the native T-domain. A representative design showed improved soluble expression, refolding, and a 12 {degrees}C higher melting temperature, while molecular dynamics simulations indicated preserved global stability but reshaped, state-dependent interdomain contact networks. Together, these results establish generative design as an effective route to context-conditioned optimization and reprogramming of biosynthetic assembly lines.

synthetic biology↗

Accelerating ligand discovery by combining Bayesian optimization with MMGBSA-based binding affinity calculations

Predicting protein-ligand binding affinity with high accuracy is critical in structure-based drug discovery. While docking methods offer computational efficiency, they often lack the precision required for reliable affinity ranking. In contrast, molecular dynamics (MD)-based approaches such as MMGBSA provide more accurate binding free energy estimates but are computationally intensive, limiting their scalability. To address this trade-off, we introduce an active learning framework that automates molecule selection for docking and MD simulations, replacing manual expert-driven decisions with a data-efficient, model-guided strategy. Our approach integrates fixed -- partly pre-trained deep learning -- molecular embeddings (MolFormer, ChemBERTa-2, and Morgan fingerprints) with adaptive regression models (e.g. Bayesian Ridge and Random Forest) to iteratively improve binding affinity predictions. We evaluate this approach retro-spectively on a new dataset of 59,356 chemically diverse compounds from ZINC-22 targeting the MCL1 protein using both AutoDock Vina and MMGBSA binding free energy scores. Our results show that incorporating MMGBSA scores into the active learning loop significantly enhances performance, recovering 79.9% of the top 1% binders in the whole dataset, compared to only 6.7% when using docking scores alone. Notably, MMGBSA exhibits a stronger correlation with experimental binding affinities than AutoDock Vina on our dataset and enables more accurate ranking of candidate compounds in a runtime efficient way. Furthermore, we demonstrate that a one-at-a-time acquisition active learning strategy consistently outperforms traditional batched acquisition, the latter achieving just 78.4% recovery with MolFormer and Bayesian Ridge. These findings underscore the potential of integrating deep learning-based molecular representations with MD-level accuracy in an active learning framework, offering a scalable and efficient path to accelerate virtual screening and improve hit identification in drug discovery.

bioinformatics↗