bioRxiv · 10.1101/2025.01.14.633076
Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity
Abstract
Predicting protein-ligand binding characteristics, such as affinity and kinetics, is critical for accelerating drug discovery. However, many existing computational methods face key limitations, including insufficient integration of comprehensive databases, inadequate representation of protein structural dynamics, and incomplete modeling of microscale protein-ligand interactions. To address these challenges, we introduce ProMoNet, a sequence-based pre-training and fine-tuning framework to enhance the prediction of protein-ligand binding characteristics. ProMoNet leverages protein and molecular foundation models to expand data coverage and enhance diversity. It also introduces a pre-training strategy based on protein-ligand binding site prediction, which bridges protein- and ligand-level representations to support downstream prediction tasks involving protein-ligand complexes. Our pre-training module effectively models microscale protein-ligand interactions and captures the dynamic nature of proteins, including binding site crypticity, without relying on 3-dimensional structural inputs. Notably, this module surpasses or matches state-of-the-art structure-based methods in identifying exposed and cryptic binding sites while maintaining high efficiency. Our fine-tuning module then efficiently transfers the pre-trained knowledge to downstream tasks such as binding affinity and binding kinetics prediction, achieving superior performance. The combination of ProMoNets strong performance and demonstrated efficiency across multiple tasks highlights its potential for broad applications in drug discovery. Scientific ContributionWe propose ProMoNet, a sequence-based pre-training and fine-tuning framework for protein-ligand binding characteristic prediction, where protein-ligand binding site prediction is introduced as a pre-training strategy to bridge independent protein- and ligand-level representations for downstream complex-level tasks. We design two dedicated modules, including a pre-training module that models microscale protein-ligand interactions and captures protein dynamics, as well as a fine-tuning module that efficiently integrates the pre-trained representations for downstream tasks. Even compared to structure-based methods, ProMoNet matches state-of-the-art performance in exposed and cryptic binding site identification and delivers superior results in binding affinity and kinetics prediction, making it a promising tool for drug discovery.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhang, S., Xie, L., Tiourine, D.. 2025-01-19. Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity. https://doi.org/10.1101/2025.01.14.633076
Cite the original work for its findings. Save a collection to share your selection of sources.