bioRxiv Science⌕ Search

Biology subjects

Dewan, S.

Publications and source records attributed to Dewan, S..

2 recordsLinked to original sources

An Improved Systematic Method for Constructing ecGEMs using a Protein-Chemical Transformer

Enzyme-constrained genome-scale metabolic models (ecGEMs) have improved Flux Balance Analysis (FBA) by incorporating enzyme turnover numbers (kcats). Since in-vivo kcat data is costly to obtain and therefore scarce, we present a novel multi-modal transformer-based approach with cross-attention to predict kcat values for Escherichia coli using enzyme amino acid sequences and SMILES annotations of reaction substrates. For heteromeric enzymes, we evaluate multiple subunit kcat aggregation strategies. We benchmark ecGEMs constructed with these strategies against current state-of-the-art models using experimental growth rates, 13C fluxes, and enzyme abundances, and prior to any calibration outperform or match existing methods. We also devise a new calibration method using flux control coefficients (derivatives of log flux with respect to log kcat), which we show to be identical to enzyme cost at the FBA optimum. Using these coefficients, we identify 8 key kcat values to recalibrate using experimental data, subsequently achieving superior performance to the current state-of-the-art with 81% fewer calibrations.

systems biology↗

Protein Language Models in Directed Evolution

The dominant paradigms for integrating machine-learning into protein engineering are de novo protein design and guided directed evolution. Guiding directed evolution requires a model of protein fitness, but most models are only evaluated in silico on datasets comprising few mutations. Due to the limited number of mutations in these datasets, it is unclear how well these models can guide directed evolution efforts. We demonstrate in vitro how zero-shot and few-shot protein language models of fitness can be used to guide two rounds of directed evolution with simulated annealing. Our few-shot simulated annealing approach recommended enzyme variants with 1.62 x improved PET degradation over 72 h period, outperforming the top engineered variant from the literature, which was 1.40 x fitter than wild-type. In the second round, 240 in vitro examples were used for training, 32 homologous sequences were used for evolutionary context and 176 variants were evaluated for improved PET degradation, achieving a hit-rate of 39 % of variants fitter than wild-type.

bioinformatics↗