bioRxiv Science⌕ Search

Biology subjects

Ghukasyan, T.

Publications and source records attributed to Ghukasyan, T..

2 recordsLinked to original sources

Overcoming the accuracy-generalization tradeoff in docking and scoring for prospective virtual screening

Virtual screening promises access to tens of billions of synthetically accessible, diverse compounds, yet it is rarely used as a primary hit-discovery strategy in contemporary drug-discovery campaigns. We argue that this gap reflects the real-world underperformance of the underlying docking and scoring methods: classical docking is generalizable but limited in accuracy by simple functional forms and parsimonious parameterization, whereas recent machine-learning approaches are highly expressive but do not generalize well to novel molecules and pockets, their reported accuracy often inflated by train-test leakage. To address these challenges, we introduce DODock and DOScore, docking and scoring ML/physics hybrid frameworks that are also highly expressive, yet generalize much better out of distribution compared with the prior ML approaches. This generalization has been prospectively tested in several ways. First, DODocks blind prediction of a drug candidate binding to PCSK9 was compared to the crystal structure that was subsequently solved, recovering the binding pose to 1.2 [A] RMSD. We also used DODock and DOScore in prospective virtual screening campaigns against four therapeutic targets, spanning an ectoenzyme (CD73), a kinase (IRAK4), an extended-substrate protease (FXI), and an allosteric inhibition of protein-protein interface (IL17). These screens yielded many chemically novel, biochemically and cellularly active inhibitors. In the case of CD73, which is a historically challenging target for virtual screening, our screen resulted in a roughly hundredfold improvement in hit rate over a recent machine-learning screen. Our results indicate that the apparent ceiling in the accuracy of virtual screening that seemed to have somewhat plateaued in the last two decades is not fundamental, and that structure-based interrogation of ultralarge chemical space may eventually become a credible primary route to novel chemical matter.

biophysics↗

SMART DATA FACTORY: VOLUNTEER COMPUTING PLATFORM FOR ACTIVE LEARNING-DRIVEN MOLECULAR DATA ACQUISITION

This paper presents the Smart Distributed Data Factory (SDDF), an AI-driven distributed computing platform designed to address challenges in drug discovery by creating comprehensive datasets of molecular conformations and their properties. SDDF uses volunteer computing, leveraging the processing power of personal computers worldwide to accelerate quantum chemistry (DFT) calculations. To tackle the vast chemical space and limited high-quality data, SDDF employs an ensemble of machine learning models to predict molecular properties and selectively choose the most challenging data points for further DFT calculations. The platform also generates new molecular conformations using molecular dynamics with the forces derived from these models. SDDF makes several contributions: the volunteer computing platform for DFT calculations; an active learning framework for constructing a dataset of molecular conformations; a large public dataset of diverse ENAMINE molecules with calculated energies; an ensemble of state-of-the-art ML models for accurate energy prediction. The energy dataset was generated to validate the SDDF approach of reducing the need for extensive calculations. With its strict scaffold split, the dataset can be used for training and benchmarking energy models. By combining active learning, distributed computing, and quantum chemistry, SDDF offers a scalable, cost-effective solution for developing accurate molecular models and ultimately accelerating drug discovery.

biophysics↗