bioRxiv Science⌕ Search

Biology subjects

Mansour, B.

Publications and source records attributed to Mansour, B..

2 recordsLinked to original sources

From Predicted Ki to Surrogate IC50: Similarity-Guided Empirical Calibration of Drug Target Affinity Predictions

Drug target affinity models return the endpoint on which they are trained, whereas medicinal chemistry decisions are often made with a different assay readout. Here, we trained a DeepPurpose model to estimate inhibition constants (Ki) from molecular graphs and protein sequences and asked whether those predictions could be aligned empirically with measured IC50 values without treating Ki and IC50 as interchangeable. A BindingDB-trained checkpoint retained useful cross-target ranking on the Davis kinase benchmark without Davis training data (concordance index 0.860). We then rebuilt the Ki training set from ChEMBL 37 records coded as Binding assays (assay_type = 'B'), after removing censored records and targets with poor replicate reproducibility. This reduced median fold error from 16.35x to 9.23x on a leakage-cleaned kinase panel and from 15.30x to 7.90x on a 14-target non-kinase panel before any IC50 calibration. Similarity-guided leave-one-out calibration further reduced the non-kinase panel median error to 3.12x for the original checkpoint and 3.22x for the ChEMBL checkpoint at Tanimoto T = 0.6. Because retraining removed a substantial part of the apparent correction, we interpret the calibration as a target- and chemistry-dependent empirical offset between model output and IC50 assay space, not as a mechanistic Ki-to-IC50 conversion. In a separate project-level stratification of public records, median replicate variability was 2.41x for the subset classified as biochemical Ki and 3.30x for biochemical IC50; a small-sample correction placed the Ki variability nearer 2.8x. These values provide an empirical scale for the remaining calibration error rather than a theoretical performance limit. Similarity, rather than the number of calibrators, governed the main accuracy coverage trade-off. The resulting values are surrogate IC50 estimates for cross-target triage; within-target ranking remains a limitation of the present architecture.

biochemistry↗

ADMET Property Prediction with Quantum-Inspired Preprocessing

Accurate prediction of Absorption, Distribution, Metabolism, Excretion, and Toxicity (ADMET) properties is a central challenge in early-stage drug discovery, where experimental determination remains costly and time-consuming. In this work, we propose a quantum-inspired preprocessing framework in which statistical dependencies among molecular descriptors are encoded into a parameterised many-body Hamiltonian, and the expectation values obtained by simulating its time evolution serve as additional inputs to a gradient-boosted ensemble model (CatBoost). Mutual information (MI) is used both to select the most informative descriptors and to set the coupling strengths of the Hamiltonian, so that the induced entanglement structure reflects empirically measured feature correlations; the evolution is realised with a short digitised-counterdiabatic schedule that generates a compact set of expectation-value features while keeping the circuit shallow. The resulting quantum-derived feature vectors are concatenated with the full MapLight descriptor set, concatenated ECFP, Avalon, and ErG fingerprints together with RDKit physicochemical properties, before training. We evaluate the pipeline on the AqSolDB aqueous solubility benchmark from the Therapeutics Data Commons (TDC) platform, achieving a mean absolute error (MAE) of 0.746 {+/-} 0.006 log(mol/L), which is within the reported error bars of the current top-performing model on the TDC leaderboard (MAE = 0.741 {+/-} 0.013). Ablation experiments show that the quantum-derived features match classical second-degree polynomial interaction features derived from the same MI-selected subset, while forming a far more compact representation (85 quantum features versus up to 4,950 polynomial terms, an approximately 58-fold reduction). SHapley Additive exPlanations (SHAP) analysis identifies the physicochemical drivers of solubility predictions, offering interpretable insight into model behaviour. These results demonstrate that MI-guided Hamiltonian feature extraction can reproduce the performance of strong classical interaction models on aqueous solubility while generating a compact, interpretable feature representation that is compatible with future quantum execution.

bioinformatics↗