bioRxiv Science⌕ Search

Biology subjects

Poelmans, R.

Publications and source records attributed to Poelmans, R..

2 recordsLinked to original sources

DESPOT: Direction-Enhanced Scoring POTentials

Knowledge-based potentials (KBPs) remain among the most reliable and interpretable scoring functions for protein-ligand interactions, yet most share two structural limitations. They assume that the space around each protein atom is isotropic, and their interaction-conditioned reference state cannot represent regions of space that are preferentially left empty. We introduce DESPOT (Direction-Enhanced Scoring POTentials), an all-atom anisotropic KBP that overcomes both. DESPOT classifies atoms into isotropic, axially symmetric, and fully anisotropic symmetry classes from their hybridization and bonding environment, and discretizes the surrounding interaction space using the according symmetry. By adopting a positionally averaged reference state and using a void ligand atom type, it learns, for every point around a protein atom, the probability that the point is occupied by a given ligand atom type or preferentially left empty - a ligand-independent description that naturally encodes steric exclusion. This occupancy-conditioned potential captures the precise, atom-level placement of ligand atoms; we pair it with a complementary geometry-conditioned, residue-level formulation (DESPOT-screen, in the spirit of KORP-PL) and combine the two inverse-Boltzmann scores into a consensus score, DESPOT-combo. Derived from 110,943 curated, energy-minimized complexes drawn from the CROWN database and evaluated on the CASF-2016 benchmark, DESPOT achieves competitive scoring power (Pearson r = 0.61), while DESPOT-combo attains best-in-class docking power (89.5% top-1 success); all anisotropic DESPOT variants significantly outperform isotropic KBPs and established empirical scoring functions in virtual screening. Anisotropy is decisive for rejecting geometrically implausible poses, and uniting the atom-level precision of DESPOT with the implicit flexibility tolerance of the residue-level score yields the most consistent performance across tasks. Because the same occupancy-conditioned potentials can be evaluated over an empty grid, DESPOT generates molecular interaction fields as well, unifying pose scoring with direction-aware binding-site characterization within a single interpretable model.

bioinformatics↗

CROWN: Curated Repository Of Well-resolved Noncovalent interactions

The development of machine-learning models for protein-ligand interactions is constrained by the quality and diversity of the available structural data. Existing resources force researchers into a trade-off: carefully curated collections such as PDBBind and HiQBind offer high structural reliability but cover only a narrow slice of the Protein Data Bank (PDB), whereas large-scale resources such as PLINDER provide broad coverage with minimal quality control. We present CROWN (Curated Repository Of Well-resolved Non-covalent interactions), a machine-learning-ready dataset that reconciles scale and rigor through a fully automated preprocessing pipeline. Starting from the PDB database, CROWN applies a series of interleaved quality filters and processing stages that address crystallographic resolution, ligand identity, pocket completeness, structural repair, interaction quality, and protonation at physiological pH. The pipeline finishes with a constrained energy-minimization step built on custom flat-bottomed restraints - a step absent from all existing protein-ligand datasets - that balances crystallographic evidence against the relaxation of intramolecular strain. By reconciling the heterogeneous refinement practices of different depositions without distorting the experimentally observed binding geometry, this step yields a structurally uniform collection of 178,263 complexes, representing a roughly four-fold increase in protein diversity over PDBBind and HiQBind. Rather than organizing the data around sparsely available, bias-prone binding affinities, CROWN adopts a geometry-centric design philosophy that treats the three-dimensional arrangement of atoms at the binding interface as a self-consistent source of information. To demonstrate its value as a training resource, we trained two knowledge-based scoring functions on CROWN and benchmarked them on CASF-2016: relative to HiQBind-trained counterparts, CROWN-trained models showed markedly improved ranking power (mean Spearman correlation rising from 0.509 to 0.637) and docking power (top-1 near-native pose recovery of 0.785 versus 0.724). Because CROWN imposes no requirement for affinity labels, it can in principle support any model that learns from or is evaluated against protein-ligand complex structures. We anticipate that it will serve as a broadly useful resource for tasks such as the training of binder generation, protein design or protein folding models conditioned on bound ligands, the development of scoring functions or benchmarking of interaction-prediction methods.

bioinformatics↗