bioRxiv Science⌕ Search

Biology subjects

Weskamp, N.

Publications and source records attributed to Weskamp, N..

2 recordsLinked to original sources

Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery

Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.

bioinformatics↗

Extending ligand efficacy indices with compound pharmacokinetic characteristics towards holistic Compound Quality Scores

The suitability of a small molecule to become an oral drug is often assessed by simple physicochemical rules, the application of ligand efficacy scores (combining physicochemical properties with potency) or by multi-parameter composite scores based on physicochemical compound properties. These rules and scores are empirical and typically lack mechanistic background, such as information on pharmacokinetics (PK). We introduce a new type of Compound Quality Scores (specifically called dose-scores and cmax-scores), which explicitly include predicted or when available experimentally determined PK parameters, such as volume of distribution, clearance and plasma protein binding. Combined with on-target potency, these scores are surrogates for an estimated dose or the corresponding cmax. These Compound Quality Scores allow for prioritization of compounds in test cascades, and by integrating machine learning based potency and PK predictions, these scores allow prioritization for synthesis. We demonstrate the complementary and in most cases the superiority to existing efficiency metrics (such as ligand efficiency scores) by project examples.

pharmacology and toxicology↗