bioRxiv Science⌕ Search

Biology subjects

Oswal, K.

Publications and source records attributed to Oswal, K..

3 recordsLinked to original sources

Kryptix-1: Conformational Ensemble Sampling Substantially Improves Cryptic Pocket Detection Relative to Static-Structure and Prior Computational Approaches

Cryptic binding pockets; sites absent or occluded in a proteins resting-state structure that become druggable only in specific, transiently populated conformations; represent one of the largest untapped opportunities in structure-based drug discovery. Large-scale structural surveys estimate that cryptic pockets occur in on the order of one in six protein families genome-wide, and that accounting for them could expand the druggable fraction of the disease-associated human proteome from roughly 40% to nearly 80%. Detecting cryptic pockets computationally has historically forced a choice btw physics-based conformational sampling, which is reliable but too computationally expensive to deploy across more than a handful of targets, and fast machine-learning pocket predictors trained on static structures, which frequently over-predict and lose the precision needed for practical triage. We report Kryptix-1, an in-house pipeline developed at Covenant Biosciences that combines conformational ensemble generation with a consensus pocket-scoring algorithm, benchmarked here against a matched static-structure baseline and against the published literature on a 10-protein sample from CryptoBench, an independently curated cryptic-pocket reference dataset. On the fields own standard residue-overlap metric (Jaccard index 0.5), Kryptix-1 succeeds on 80% of benchmark proteins, compared to 20% for the static baseline and approximately 40-45% for the best-performing methods reported in an independent comparative evaluation using the same metric. We report these results alongside an explicit accounting of where Kryptix-1s advantage is smaller or reverses, and we are explicit throughout that this is a small pilot evaluation, not a fully powered validation study.

bioinformatics↗

Mavchen 1: A Conformational Ensemble Platform for Protein Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

Deep learning structure predictors, most prominently AlphaFold2 (the field-standard tool benchmarked against throughout this study), have substantially expanded access to protein structural information, yet characteristically return a single static conformation per target. This is an incomplete representation of the binding-competent state for the many pharmacologically relevant targets whose recognition geometry is intrinsically dependent on receptor flexibility, including cryptic-pocket, induced-fit, and water-mediated binding mechanisms. We present a category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein-ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes. Considering the most accurate pose available from each methods full candidate output, ensemble-derived poses achieved lower RMSD to the experimental structure than AlphaFold on 21 of 29 targets (72.4%), with a mean RMSD of 3.39 [A] versus 5.60 [A]: a clear, statistically decisive advantage (paired Wilcoxon signed-rank test, W = 93.0, p = 0.0060). Rather than being diffuse, this advantage was concentrated precisely where mechanistic theory predicts it should be: in induced-fit and water-mediated categories, the classes in which static-structure prediction is expected to be least representative of the bound state: a result that constitutes direct, quantitative confirmation of the ensemble hypothesis, not merely a favorable average. Independent assessment against a field-standard physical-validity framework confirmed that this accuracy gain was achieved without any trade-off in chemical or geometric realism. We further quantify, rather than assume, the extent to which this advantage is recoverable by fully autonomous pose selection, using a proprietary ensemble-aware scoring model with no access to the correct answer, and report a substantial, discriminative signal (cross-validated mean AUC 0.92) with a partial, and clearly characterized, recovery under the strictest accuracy criteria (mean AUPR 0.36), which we identify as the principal, now precisely quantified, determinant of near-term translational progress. Under this same fully autonomous, ground-truth-blind setting, AlphaFolds own top-ranked poses currently match or modestly exceed Mavchen-1s autonomously selected poses on strict success-rate criteria (e.g., 17.2% vs. 20.7% at the combined RMSD-and-validity threshold), a result we report without qualification as the clearest current benchmark for near-term development. Together, these results provide compelling, statistically rigorous evidence that conformational ensemble sampling is a mechanistically grounded and substantial source of improved pose accuracy relative to static-structure prediction, and establish a quantitative benchmark against which continued methodological development can be measured and demonstrably improved upon.

bioinformatics↗

CHIMIYA-1: An Autoselection Foundation Model for ADMET Property Prediction, Rigorously Benchmarked Against the Therapeutics Data Commons ADMET Group

Accurate, generalizable prediction of absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties remains one of the highest-leverage unsolved problems in computational drug discovery, and late-stage attrition driven by ADMET liabilities continues to be a dominant cost driver in pharmaceutical research and development. The Therapeutics Data Commons (TDC) ADMET Group has emerged as the fields most widely adopted public benchmark, comprising 22 endpoints under standardized scaffold-split evaluation. In this work we report a comprehensive evaluation of CHIMIYA-1, a proprietary autoselection foundation model developed by Covenant Biosciences, against the full TDC ADMET Group. Departing from common practice in the field, every reported score is the mean and standard deviation of five independently seeded end-to-end evaluation runs (TDCs own minimum submission standard, which we find is not met by all public leaderboard entries), and all 22 endpoints were additionally subjected to an explicit train/test structural-overlap audit prior to reporting, finding zero overlaps on any endpoint. Despite this deliberately conservative evaluation standard, CHIMIYA-1 ranks first among all publicly listed methods on four endpoints, places within the top decile of the field on twenty of twenty-two endpoints (91%), and attains a mean percentile standing near the 74th percentile across the full benchmark, with particular strength on toxicity and physicochemical-property endpoints. We further show that several top-ranked public comparators on this benchmark have been independently found to exhibit confirmed data leakage, a finding that, if anything, understates CHIMIYA-1s relative standing. All results were obtained on commodity single-GPU workstation hardware without recourse to distributed or cloud-scale training infrastructure. We discuss these results in the context of benchmark reporting norms in molecular machine learning and outline ongoing extensions, including continuous prospective-data retraining and CUDA-level throughput optimization of the underlying selection pipeline.

bioinformatics↗