bioRxiv Science⌕ Search

Biology subjects

Kroll, A.

Publications and source records attributed to Kroll, A..

4 recordsLinked to original sources

Machine learning models for the prediction of enzyme properties should be tested on proteins not used for model training

The recently published DLKcat model, a deep learning approach for predicting enzyme turnover numbers (kcat), claims to enable high-throughput kcat predictions for metabolic enzymes from any organism and to capture kcat changes for mutated enzymes. Here, we critically evaluate these claims. We show that DLKcat predictions become positively misleading for enzymes with less than 60% sequence identity to the training data, performing worse than simply assuming a mean kcat value for all reactions. Furthermore, DLKcats ability to predict mutation effects is much weaker than implied, capturing only 3% of the experimentally observed variation across mutants not included in the training data. These findings highlight significant limitations in DLKcats generalizability and its practical utility for predicting kcat values for novel enzyme families or mutants, which are crucial applications in fields such as metabolic modeling.

bioinformatics↗

Turnover number predictions for kinetically uncharacterized enzymes using machine and deep learning

The turnover number kcat, a measure of enzyme efficiency, is central to understanding cellular physiology and resource allocation. As experimental kcat estimates are unavailable for the vast majority of enzymatic reactions, the development of accurate computational prediction methods is highly desirable. However, existing machine learning models are limited to a single, well-studied organism, or they provide inaccurate predictions except for enzymes that are highly similar to proteins in the training set. Here, we present TurNuP, a general and organism-independent model that successfully predicts turnover numbers for natural reactions of wild-type enzymes. We constructed model inputs by representing complete chemical reactions through difference fingerprints and by representing enzymes through a modified and re-trained Transformer Network model for protein sequences. TurNuP outperforms previous models and generalizes well even to enzymes that are not similar to proteins in the training set. Parameterizing metabolic models with TurNuP-predicted kcat values leads to improved proteome allocation predictions. To provide a powerful and convenient tool for the study of molecular biochemistry and physiology, we implemented a TurNuP web server at https://turnup.cs.hhu.de.

bioinformatics↗

The substrate scopes of enzymes: a general prediction model based on machine and deep learning

For a comprehensive understanding of metabolism, it is necessary to know all potential substrates for each enzyme encoded in an organisms genome. However, for most proteins annotated as enzymes, it is unknown which primary and/or secondary reactions they catalyze [1], as experimental characterizations are time-consuming and costly. Machine learning predictions could provide an efficient alternative, but are hampered by a lack of information regarding enzyme non-substrates, as available training data comprises mainly positive examples. Here, we present ESP, a general machine learning model for the prediction of enzyme-substrate pairs, with an accuracy of over 90% on independent and diverse test data. This accuracy was achieved by representing enzymes through a modified transformer model [2] with a trained, task-specific token, and by augmenting the positive training data by randomly sampling small molecules and assigning them as non-substrates. ESP can be applied successfully across widely different enzymes and a broad range of metabolites. It outperforms recently published models designed for individual, well-studied enzyme families, which use much more detailed input data [3, 4]. We implemented a user-friendly web server to predict the substrate scope of arbitrary enzymes, which may support not only basic science, but also the development of pharmaceuticals and bioengineering processes.

bioinformatics↗

Prediction of Michaelis constants from structuralfeatures using deep learning

The Michaelis constant KM describes the affinity of an enzyme for a specific substrate, and is a central parameter in studies of enzyme kinetics and cellular physiology. As measurements of KM are often difficult and time-consuming, experimental estimates exist for only a minority of enzyme-substrate combinations even in model organisms. Here, we build and train an organism-independent model that successfully predicts KM values for natural enzyme-substrate combinations using machine and deep learning methods. Predictions are based on a task-specific molecular fingerprint of the substrate, generated using a graph neural network, and the domain structure of the enzyme. Model predictions can be used to estimate enzyme efficiencies, to relate metabolite concentrations to cellular physiology, and to fill gaps in the parameterization of kinetic models of cellular metabolism.

systems biology↗