bioRxiv Science⌕ Search

Biology subjects

Boorla, V.

Publications and source records attributed to Boorla, V..

2 recordsLinked to original sources

Accessing Enzyme Kinetic Data and Prediction Methods at Scale

Enzyme kinetic parameters inform metabolic models, yet experimental measurements are sparse. A growing body of work predicts them from protein and substrate features, but software fragmentation hinders adoption, so downstream tools lock into the most accessible method. We present OpenKinetics Predictor (at predictor.openkinetics.org), an open-source platform integrating thirteen methods in isolated environments behind one interface. The platform optionally reports similarity between query proteins and each method's training data to contextualise reliability. A common featurisation-prediction abstraction keeps it extensible, and independent parties, including original authors, contributed many methods. We pair it with a data portal (at data.openkinetics.org) that exposes CatLog, a curated kinetic dataset, with precomputed embeddings, predicted binding sites, and standardised splits. Both offer a web interface and an API, and the GECKO modelling toolbox already calls the predictor API. As a case study, we predict across an E. coli model and find inter-predictor agreement varies with metabolic context and data availability.

systems biology↗

BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings

MotivationReliable predictions of protein-protein binding affinities are essential for molecular biology and therapeutic discovery. However, most computational methods rely on three-dimensional structural models, which are often unavailable for many complexes. ResultsWe introduce BindPred, a structure-agnostic input framework that predicts affinities directly from amino acid sequences by combining embeddings from large protein language models with gradient boosting trees. On the PPB-Affinity benchmark, which comprises 11,919 diverse complexes, BindPred achieves a Pearson correlation coefficient of 0.86 in random split five-fold cross-validation, where the training and test sets share <30% global sequence identity. Ablation analysis indicates that evolutionary embeddings alone capture most of the predictive signals, while augmenting with physics-based energy terms from PyRosetta and BindCraft increases the correlation by only 0.01. A more stringent protein-level split that places entire protein families (wild-type and all mutants) exclusively in either training or testing sets, results in only a modest decline in performance, demonstrating robust generalization to novel interaction pairs. Because BindPred operates exclusively on sequence input, it enables rapid inference (approximately 3 million complexes per GPU (T4) hour), making proteome-scale screening computationally feasible. AvailabilityThe pretrained model and inference pipeline are available in a Google Colab notebook: BindPred Colab notebook. The training dataset, code, and model weights are available on the Hugging Face: BindPred Contactcostas@psu.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗