bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.07.26.740840

Mavchen 1: A Conformational Ensemble Platform for Protein Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

Abstract

Deep learning structure predictors, most prominently AlphaFold2 (the field-standard tool benchmarked against throughout this study), have substantially expanded access to protein structural information, yet characteristically return a single static conformation per target. This is an incomplete representation of the binding-competent state for the many pharmacologically relevant targets whose recognition geometry is intrinsically dependent on receptor flexibility, including cryptic-pocket, induced-fit, and water-mediated binding mechanisms. We present a category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein-ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes. Considering the most accurate pose available from each methods full candidate output, ensemble-derived poses achieved lower RMSD to the experimental structure than AlphaFold on 21 of 29 targets (72.4%), with a mean RMSD of 3.39 [A] versus 5.60 [A]: a clear, statistically decisive advantage (paired Wilcoxon signed-rank test, W = 93.0, p = 0.0060). Rather than being diffuse, this advantage was concentrated precisely where mechanistic theory predicts it should be: in induced-fit and water-mediated categories, the classes in which static-structure prediction is expected to be least representative of the bound state: a result that constitutes direct, quantitative confirmation of the ensemble hypothesis, not merely a favorable average. Independent assessment against a field-standard physical-validity framework confirmed that this accuracy gain was achieved without any trade-off in chemical or geometric realism. We further quantify, rather than assume, the extent to which this advantage is recoverable by fully autonomous pose selection, using a proprietary ensemble-aware scoring model with no access to the correct answer, and report a substantial, discriminative signal (cross-validated mean AUC 0.92) with a partial, and clearly characterized, recovery under the strictest accuracy criteria (mean AUPR 0.36), which we identify as the principal, now precisely quantified, determinant of near-term translational progress. Under this same fully autonomous, ground-truth-blind setting, AlphaFolds own top-ranked poses currently match or modestly exceed Mavchen-1s autonomously selected poses on strict success-rate criteria (e.g., 17.2% vs. 20.7% at the combined RMSD-and-validity threshold), a result we report without qualification as the clearest current benchmark for near-term development. Together, these results provide compelling, statistically rigorous evidence that conformational ensemble sampling is a mechanistically grounded and substantial source of improved pose accuracy relative to static-structure prediction, and establish a quantitative benchmark against which continued methodological development can be measured and demonstrably improved upon.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Varghese, R., Tiwary, P., Oswal, K.. 2026-07-29. Mavchen 1: A Conformational Ensemble Platform for Protein Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark. https://doi.org/10.64898/2026.07.26.740840

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A meta-interaction basis for cell-cell communication in tissues

Tissue function depends on signals exchanged between cells and the responses they elicit. Yet whether diverse cell-cell interactions in situ form recurrent sender-receiver programs remains unclear. We present SpiderNet, an interpretable representation-learning framework that discovers such directed programs as a compact basis of cell-cell meta-interactions (MIs) from spatial transcriptomics. SpiderNet jointly learns which sender regulators, ligand-receptor pairs, and receiver targets define each MI and where each program is active across neighboring cell pairs. The resulting representation traces multicellular relays and links communication to cell states, perturbation responses, and phenotypes. SpiderNet recovers ground-truth MIs and their molecular components in simulations and, in real tissues, shows stronger direction-specific agreement with independently curated regulatory programs in senders and receivers than alternative methods. Across more than 5.8 million spatially profiled cells, SpiderNet resolves an SPP1-THBS relay linking monocytes, fibroblasts, and tumor cells within an immune-suppressive ovarian cancer niche, predicts T-cell responses to held-out melanoma-cell perturbations, and identifies a T-cell-associated brain-aging program and age-predictive signals that transfer across regions and platforms. It reveals a recurrent pan-cancer COLLAGEN-linked fibroblast-tumor program whose projected abundance in independent cohorts is associated with poorer survival and non-response to immunotherapy. SpiderNet thus establishes MIs as a reusable organizational layer between molecular interactions and tissue phenotypes, providing a framework to resolve, compare, trace, and perturb multicellular regulation in situ.

bioinformatics↗

Heterogeneous Graph Contrastive Learning for Drug-Gene-Disease Motif Prediction

Drug repurposing and target discovery offer critical strategies for advancing therapeutic development by uncovering the potential biological pathways and novel associations among drugs, genes, and diseases. However, experimental discovery remains expensive and time-consuming, which limits the scalability of large-scale studies. In addition, existing computational approaches often struggle to effectively integrate heterogeneous biomedical data, capture the complex higher-order topological signatures of biological interactomes, and generalize to unseen entities. Here, we present HANAMI (Heterogeneous grAph coNtrastive leArning for drug-gene-disease Motif predIction), a multi-view deep graph learning framework designed to model complex interactions among drugs, genes, and diseases. HANAMI integrates diverse heterogeneous biomedical knowledge, including chemical structures, genomic sequences, and clinical phenotypes, and leverages relation-aware topology encoding, structure-aware aggregation, and contrastive learning to enable accurate motif prediction with biological context from the network. Systematic evaluation on benchmark datasets shows that HANAMI achieves up to 6% improvements over existing state-of-the-art methods in predicting drug-gene-disease motifs. The framework further demonstrates strong inductive generalization, maintaining an [~]18% performance advantage in zero-shot settings involving previously unseen entities. Beyond predictive performance, HANAMI effectively prioritizes drug-disease relationships investigated in Phase II or III trials while identifying candidate genes that suggest plausible mechanistic links. Together, HANAMI provides a computational framework for interpreting complex biomedical interactions, offering a scalable foundation to accelerate drug repurposing and therapeutic innovation.

bioinformatics↗

PTMExplorer: A Multi-Dimensional Integrative Visualization Platform for Protein Post-Translational Modification Function and Structure

Deciphering the functions of post-translational modifications (PTMs) is a critical bridge connecting large-scale modification proteomics data to mechanistic studies. However, most existing tools for visualizing PTM omics data are limited to site catalogs or single-dimensional feature displays. They lack the capability to simultaneously map user-derived differential modification sites onto multi-dimensional contexts, including protein three-dimensional (3D) structure, evolutionary conservation, functional sites, and disease associations. This limitation makes it difficult for researchers to rapidly assess the biological importance of candidate sites from among a vast number of differentially modified sites. Here, we present PTMExplorer, an interactive platform for the multi-dimensional visualization of protein PTMs. PTMExplorer comprises three core modules: PTM Inspector, built upon ProtVista, provides a multi-track, sequence-feature integrated view incorporating intrinsically disordered region (IDR) prediction (via flDPnn), surface accessibility calculation (via FreeSASA), and UniProt functional annotations; PTM 3D Locator, leveraging the Nightingale/Mol* engine, anchors modification sites onto AlphaFold/Protein Data Bank (PDB) 3D structures through residue mapping via PDBe-SIFTS; and PTM Overview, utilizing the R circlize package, presents a panoramic polar circos plot illustrating modification distribution and inter-group differential regulation. Additionally, three major disease-associated modification databases (PTMD, qPTM, and PhosCancer) are integrated as PTM-Disease Nexus, enabling co-localization comparison between user-defined differential sites and reported disease-related sites. PTMExplorer currently supports eight model organisms, accepts user-uploaded differential analysis results, and provides multi-dimensional annotations and various visualization options (https://www.bioladder.cn/PTMExplorer/). Using a multi-omics dataset from hepatocellular carcinoma (18 patients, 9 modification types) as a case study, we demonstrate the practical utility of PTMExplorer in screening potential biomarkers, revealing multi-modification coordination mechanisms, and distinguishing between absolute and relative quantification patterns.

bioinformatics↗