bioRxiv Science⌕ Search

Biology subjects

Suutari, B.

Publications and source records attributed to Suutari, B..

2 recordsLinked to original sources

Learning Binding Affinities via Fine-tuning of Protein and Ligand Language Models

Accurate in-silico prediction of protein-ligand binding affinity is essential for efficient hit identification in large molecular libraries. Commonly used structure-based methods such as docking often fail to rank compounds effectively, and free energy-based approaches, while accurate, are too computationally intensive for large-scale screening. Existing deep learning models struggle to generalize to new targets or drugs, and current evaluation methods often do not accurately reflect real-world performance. We introduce BALM, a deep learning framework that predicts binding affinity using pre-trained protein and ligand language models. We also propose improved evaluation strategies with diverse data sets and metrics to assess model performance to new targets better. Using the BindingDB dataset, BALM generalises unseen drugs, scaffolds, and targets. In few-shot scenarios for targets such as USP7 and Mpro, it outperforms traditional machine learning and docking methods, including AutoDock Vina. Adoption of our target-based evaluation methods will allow a more stringent evaluation of machine learning-based scoring tools. Our protein prediction framework shows good performance, is computationally efficient, and is highly adaptable within this evaluation setting, making it practical for early-stage drug discovery screening.

bioinformatics↗

Benchmarking active learning protocols for ligand binding affinity prediction

Active learning (AL) has become a powerful tool in computational drug discovery, enabling the identification of top binders from vast molecular libraries with reduced costs for relative binding free energy calculations and experiments. To design a robust AL protocol, it is important to understand the influence of AL parameters, as well as the features of the datasets on the outcomes. We use four affinity datasets for different targets (TYK2, USP7, D2R, Mpro) to systematically evaluate the performance of machine learning models (Gaussian Process model, Chemprop), sample selection protocols, as well as the batch size based on metrics describing the overall predictive power of the model (R2, Spearman rank, RMSE) as well as the accurate identification of top 2% / 5% binders (Recall, F1 score). Both models have a comparable Recall of top binders on large datasets, but the GP models surpass Chemprop when training data is sparse. A larger initial batch size, especially on diverse datasets, increased the Recall of both models as well as overall correlation metrics. However, for subsequent cycles, smaller batch sizes of 20 or 30 compounds proved to be desirable. Furthermore, the presence of Gaussian noise to the data, up to a certain threshold, still allowed the model to identify clusters with top-scoring compounds. However, excessive noise (<1{sigma}) did impact the models predictive and exploitative capabilities. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/568570v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@cefb11org.highwire.dtl.DTLVardef@c522bcorg.highwire.dtl.DTLVardef@6b8f18org.highwire.dtl.DTLVardef@17f7706_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOTOC GraphicC_FLOATNO C_FIG

bioinformatics↗