bioRxiv Science⌕ Search

Biology subjects

Shome, R.

Publications and source records attributed to Shome, R..

2 recordsLinked to original sources

Chemical Dice Integrator (CDI): A Scalable Framework for Multimodal Molecular Representation Learning

The machine learning landscape for molecular property prediction is fragmented, with numerous Featurizers each capturing a narrow, specialized view of chemical structure. This heterogeneity forces a suboptimal choice of representation a priori, limiting model generalizability. We introduce the Chemical Dice Integrator (CDI), a hierarchical framework that unifies six orthogonal molecular representations, physicochemical (Mordred), topological (GROVER), visual (ImageMol), biological (Signaturizer), quantum-mechanical (MOPAC), and linguistic (ChemBERTa), into a single, coherent embedding. The framework consists of CDI-Basic, a two-tiered autoencoder that fuses these modalities, and CDI-Generalised, a Mamba State-Space Model (SSM) that learns a direct, efficient map from SMILES strings to the unified embedding space. Extensive benchmarking across 23 classification (171 tasks) and 10 regression datasets demonstrates that CDI embeddings consistently achieve superior predictive performance compared to individual Featurizers and standard feature aggregation methods. The CDI-Generalised model achieves this performance with exceptional computational efficiency, outperforming deep learning Featurizers in terms of speed and resource overhead. Furthermore, we demonstrate that the CDI embedding is chemically intuitive, allowing for the sensitive distinction of nuanced structural variants, such as chiral enantiomers and kekulized SMILES forms. By bridging multimodal chemical intelligence with scalable, sequence-based inference, CDI offers a strong foundation for molecular machine learning.

bioinformatics↗

SynGlue: AI-Driven Designer for Clinically Actionable Multi-Target Therapeutics

The rational design of protein degraders, such as proteolysis-targeting chimeras (PROTACs), requires the simultaneous optimization of multiple molecular properties, a complex challenge that limits efficient discovery. Here, we introduce SynGlue, a generative artificial intelligence (AI) framework that addresses this challenge through two core modules: data-driven, leveraging large-scale protein-ligand intelligence, and structure-guided, for physics-aware molecular design. SynGlue harness MagnetDB, a curated database of 6.37 million experimental protein-ligand interactions, and couples it with deep learning models that quantitatively predict degradation potency (DC50), maximal degradation (Dmax), and guide ternary-complex-compatible linker design. Benchmarked against 6,935 compounds, SynGlue demonstrates superior performance in relevant pharmacology prediction. To validate SynGlue, we engineered degraders for BRD4 and GSPT1. Our data-driven design for BRD4 yielded compounds with novel warhead scaffolds (<50% warhead similarity with known PROTACs), which proved to be potent degraders in vitro (DC50 = 0.19 nM) and efficacious in vivo in mouse models. Independently, our structure-guided de novo design for GSPT1 produced ultrapotent degraders (DC50 {approx} 0.0011 M) that are also effective both in vitro and in vivo, uncovering a new oncogenic dependency. By unifying data-driven and physics-aware design, SynGlue establishes a generalizable AI framework for the rapid development of clinically relevant protein degraders, with principled extension to other multi-target modalities.

bioinformatics↗