bioRxiv Science⌕ Search

Biology subjects

Nafiiev, A.

Publications and source records attributed to Nafiiev, A..

3 recordsLinked to original sources

Sampling and ranking of protein conformations using machine learning techniques do not improve quality of rigid protein-protein docking

Rigid docking remains the most popular method of predicting protein-protein interactions in cases when experimental 3D structures of the complexes are not available. The docking often relies on known unbound (Apo) protein structures, which may differ significantly from their bound (Holo) forms. Modern machine learning (ML) based conformational sampling techniques allow generating ensembles of functionally relevant protein structures, which may be closer to their Holo forms and thus could improve the outcomes of the classical rigid protein-protein docking. Here, we sampled conformations of the protein subunits in 30 complexes from the novel PINDER dataset with two state-of-the-art ML-based techniques and evaluated their docking performance using several physics-based, data-based, and ML-based scoring functions. We showed that such conformational sampling rarely produces structures that are closer to the Holo conformations than the corresponding Apo ones. Moreover, even when such conformations are generated, none of the tested scoring functions were able to prioritize and rank them correctly. Our work highlights critical limitations in the current ML-enhanced rigid protein-protein docking workflows and emphasizes the need for new approaches that can better utilize the potential of modern techniques for conformational generation and scoring.

biophysics↗

Leveraging Large Language Models for Literature-Driven Prioritization of Protein Binding Pockets

We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with Large Language Models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLMs performance in paper relevance assessment and pocket extraction. The extracted information is then mapped onto protein structures and used to filter and merge the geometry-based predictions, generating a refined volumetric representation of biologically relevant pockets. This hybrid pipeline offers an efficient, accurate and automated method for identifying functional binding pockets, addressing a significant bottleneck in the high-throughput drug discovery workflows. The developed benchmark dataset and methodology are freely available at https://github.com/MelnychenkoM/LLM-benchmark-dataset.

biophysics↗

ArtiDock: fast and accurate machine learning approach to protein-ligand docking based on multimodal data augmentation

Classical protein-ligand docking has been a cornerstone technique in computational drug discovery for decades, but has reached an accuracy and performance plateau. Recently introduced Machine Learning (ML) based docking methods offer a promising paradigm shift, but their practical adoption is hampered by accuracy-to-speed trade-offs, inadequate benchmarking standards, and questionable chemical validity of predicted poses. In this study, we introduce ArtiDock - an ML-based docking technique optimized for high-throughput virtual screening applications. To evaluate ArtiDock, we developed a dedicated performance and accuracy benchmark for pocket-specific rigid protein-ligand docking, which mimics realistic industrial drug discovery scenarios and is based on the novel PLINDER dataset. We demonstrate that ArtiDock is 29-38% more accurate in comparison to leading open-source and commercial classical docking techniques such as AutoDock, Vina, and Glide, while providing a low computational cost. ArtiDock notably excels in challenging docking scenarios involving unbound protein structures and binding sites containing ions and structured water molecules. Our results show that ArtiDock could be considered as a method of choice in high-throughput virtual screening scenarios.

bioinformatics↗