bioRxiv Science⌕ Search

Biology subjects

Stockinger, P.

Publications and source records attributed to Stockinger, P..

3 recordsLinked to original sources

How Bias Shapes the Leaderboard: Scoring Function Performance Under Scrutiny

Structure-based scoring functions leveraging machine learning have recently demonstrated superior performance over classical scoring functions, particularly on virtual screening benchmarks. However, due to the fundamental differences between their underlying model principles and architectures, it remains unclear to what extent performance stems from an understanding of molecular binding or from exploitation of systemic biases. Thus, disentangling the factors underlying benchmark performance is essential for determining whether a scoring function will generalize to novel chemical space and succeed in prospective drug discovery. To address this need, we present a case study investigating the nature and impact of systemic biases on benchmark comparisons between different scoring function paradigms. By systematically analyzing the evaluation workflows of prominent models, we reveal pocket bias, a form of spatial coordinate frame leakage arising from static binding pocket extraction, which artificially inflates benchmark performance. To progressively eliminate these sources of bias, we benchmarked two selected graph neural network scoring functions against two minmalist machine learning models and a classical scoring function under four increasingly stringent evaluation levels, successively removing pocket bias, reducing structural data leakage, and finally evaluating on out-of-distribution (OOD) protein targets. Upon removal of pocket bias and structural data leakage, the performance of all machine learning models dropped substantially. When evaluated on out-of-distribution protein families, the classical baseline AutoDock Vina outperformed the machine learning models in five of seven virtual screening tasks and dominated the docking power evaluation. Our findings indicate that benchmark performance can be heavily shaped by evaluation design and dataset artifacts, potentially overshadowing algorithmic improvements. While the tested machine learning models remain heavily dependent on encountering familiar data distributions to achieve competitive results, AutoDock Vina demonstrated superior generalization capacity on OOD targets. This work underscores the critical need for rigorous, artifact-free benchmarking protocols to guide the development of truly prospective machine learning models for virtual screening.

bioinformatics↗

GEMS: A Generalizable GNN Framework For Protein-Ligand Binding Affinity Prediction Through Robust Data Filtering and Language Model Integration

The field of computational drug design requires accurate scoring functions to predict binding affinities for protein-ligand interactions. However, train-test data leakage between the PDBbind database and the CASF benchmark datasets has significantly inflated the performance metrics of currently available deep-learning-based binding affinity prediction models, leading to overestimation of their generalization capabilities. We address this issue by proposing PDBbind CleanSplit, a training dataset curated by a novel structure-based filtering algorithm that eliminates train-test data leakage as well as redundancies within the training set. Retraining current top-performing models on CleanSplit caused their benchmark performance to drop significantly, indicating that the performance of existing models is largely driven by data leakage. In contrast, our graph neural network model for efficient molecular scoring (GEMS) maintains high benchmark performance when trained on CleanSplit. Leveraging a sparse graph modeling of protein-ligand interactions and transfer learning from language models, GEMS is able to generalize to strictly independent test datasets.

bioinformatics↗

Chromatin Compaction Follows a Power Law Scaling with Cell Size from Interphase Through Mitosis

Coordination of mitotic chromosome compaction with cell size is crucial for proper genome segregation during mitosis. During development, DNA content remains constant but cell size evolves, necessitating a mechanism that scales chromosome compaction with cell size. In this study, we examined chromatin compaction in the developing Drosophila nervous system by analyzing the large neuronal stem cells and their smaller progeny, the ganglion mother cells. Using super-resolution 3D Stochastic Optical Reconstruction Microscopy and quantitative time-lapse fluorescence microscopy, we observed that nanoscale chromatin density during interphase scales with nuclear volume according to a power law. This scaling relationship is disrupted by inhibiting histone deacetylase activity, indicating that molecular cues rather than mechanical constraints primarily regulate chromatin compaction. Notably, this power law dependency is maintained into mitosis but the scaling exponent decreases, suggesting phase separation of chromatin events during mitotic compaction. We propose that the scaling of mitotic chromosome size relative to cell size depends on the power law behaviour of interphase chromatin volume, and that scaling of mitotic chromatin compaction is an emergent property of linear polymers undergoing phase separation with their solvent. Statement of SignificanceUnderstanding how chromatin compaction changes with cell size is essential for understanding the mechanisms ensuring accurate genome segregation during cell division. In this study, we combine Voronoi tessellation analysis of Stochastic Optical Reconstruction Microscopy and image processing methods based on fluorescence intensity spectrum analysis of live confocal images to measure chromatin volume in interphase and mitosis in cells of different sizes. This work reveals that chromatin density in Drosophila neuronal cells follows a power law relationship with nuclear volume from interphase through mitosis, suggesting a fundamental principle underlying chromosome scaling across different cell sizes. We show that this scaling is regulated by molecular cues, such as histone deacetylase activity, and is conserved through a potential phase separation during mitosis, providing new insights into the biophysical processes that govern chromatin organization. This work broadens our understanding of chromosome biology and will have implications for understanding size-dependent chromatin dynamics in other organisms.

cell biology↗