bioRxiv Science⌕ Search

Biology subjects

Ran, Z.

Publications and source records attributed to Ran, Z..

10 recordsLinked to original sources

MolX: A Geometric Foundation Model for Protein-Ligand Modelling

Understanding how small molecules interact with protein binding pockets is central to structure-based drug discovery. Accurately modelling these interactions requires capturing the 3D geometry and physicochemical complementarity of binding interfaces, yet existing computational approaches encode proteins and ligands separately or rely on simplified structural representations that do not explicitly model cross-entity spatial relationships. Such decoupled representations restrict their capacity to capture interface-level geometric constraints that arise from protein-ligand coorganisation. Here we present MolX, a Graph Transformer foundation model that jointly learns geometric and chemical representations of protein pockets and ligands from large-scale 3D structural data. Integrating over 3 million protein pockets and 5 million molecules, MolX represents both entities as E(3)-equivariant graphs to preserve spatial geometry and chemical context. The architecture employs dual E(3)-equivariant graph Transformer encoders to model pocket and ligand embeddings, ensuring representations remain invariant to rotation, translation, and reflection. MolX is pretrained using a hybrid learning paradigm that combines supervised biochemical objectives, logP and energy-gap regression, with self-supervised geometric objectives, coordinate re-construction, and atom-type prediction, fostering generalisable molecular understanding. Across eight downstream benchmarks, including antibody-drug conjugates (ADC), proteolysis-targeting chimeras (PROTAC), molecular glue, and PCBA activity prediction, as well as binding affinity and physicochemical property regression, MolX achieves consistent state-of-the-art performance and strong cross-domain generalisation. Furthermore, MolX incorporates a sparse autoencoder module to decompose latent representations into interpretable biological components, thereby revealing the pocket-ligand interactions that drive prediction outcomes. Together, MolX establishes a scalable and interpretable foundation model for molecular representation learning, providing a unified framework for predicting and interpreting complex small-molecule-protein interactions.

bioinformatics↗

Graph-based RNA structural representation reveals determinants of subcellular localization

RNA subcellular localization is a key determinant of RNA function and regulation, yet existing computational approaches rely primarily on sequence or simplified structural descriptors, limiting their scalability to long transcripts, their ability to model inter-label dependencies, and their applicability across RNA types. Here, we present GRASP, a unified graph neural network framework for predicting RNA subcellular localization using a heterogeneous graph representation that is RNA substructure-aware. GRASP presents each RNA as a multi-scale graph comprising nucleotide nodes and secondary-structure-derived substructure nodes, connected by relational edges, enabling joint modeling of base-level interactions and regional structural context. The model further incorporates multi-label dependency learning to capture co-localization patterns across cellular compartments within a unified framework. Across multiple benchmark datasets and RNA types, GRASP consistently outperforms state-of-the-art sequence-based and structure-informed methods, achieving substantial improvements in accuracy, F1 score, and AUC while maintaining strong scalability to long transcripts. In addition, the graph-based representation provides biologically interpretable insights into structural determinants of RNA localization. The source code and data are available at https://github.com/ABILiLab/GRASP, and the web server is accessible at http://grasp.biotools.bio.

bioinformatics↗

DOMINO: diffusion-optimised graph learning identifies domain structures with enhanced accuracy and scalability

Spatial transcriptomics enables in situ molecular profiling, allowing to measure the cellular transcriptional output within the tissue. As the tissue architecture is conserved, spatial domains with specific transcriptional patterns can be identified, facilitating the discovery and understanding of functional tissue compartments. Thus, several methods to uncover and identify these spatial domains have been developed. However, most of these existing methods do not scale to rapidly increasing data sizes and focus only on local structure while missing the global view of the tissue. Here, we present DOMINO, a diffusion-optimised contrastive learning framework for spatial domain detection. DOMINO utilises graph diffusion convolution to propagate information beyond immediate neighbours and jointly optimises local and, importantly, global graph structure via contrastive learning. This novel framework yields biologically interpretable domains with clearer boundaries and scales to large datasets, outperforming state-of-the-art methods across healthy and malignant benchmark datasets. We apply DOMINO to a newly generated spatial transcriptomic dataset of endometriosis-associated ovarian cancers, which could not be processed by existing domain detection methods owing to its size. We uncovered conserved proliferative and non-proliferative tumour states that recurred across these tumours and were independently validated in an external clear cell ovarian cancer spatial transcriptomic dataset. Proliferative domains were characterised by elevated expression of EIF4A1 and HSPA8, increased cell cycle activity, reduced mast cell abundance, and coordinated stromal remodelling, including altered fibroblast states and spatial organisation. In parallel, integrative analysis across tumours revealed subtype-specific multicellular ecosystems associated with either endometrioid or clear cell ovarian carcinomas, together with a tumour-excluded stromal domain that could only be resolved through the integration of spatial and transcriptional information. These findings demonstrate how well DOMINO scales up and that it uncovers biologically meaningful spatial programs spanning tumour intrinsic states, tumour microenvironment interactions, and subtype-specific tissue architecture that are not recovered by conventional expression-based clustering approaches.

bioinformatics↗

FOSL2 Directly Regulates FSHR and CYP11A1 Transcription: An Essential Transcription Factor for Gonadotropin-Dependent Folliculogenesis

The meticulous orchestration of gonadotropin-dependent folliculogenesis constitutes the cornerstone of female reproductive cyclicity and fertility, with FSH/FSHR signaling recognized as the master regulator. Achieving the necessary amplification of this signaling is essential for successful GTH-dependent folliculogenesis, yet the mechanisms remain inadequately defined. Our study utilizes single-cell and spatial transcriptomics to identify FOSL2 as an FSH-inducible transcription factor, exhibiting precise spatiotemporal co-expression with FSHR. FOSL2 knockdown in vitro resulted in notable reductions in granulosa cell proliferation, induced apoptosis, and disrupted gonadotropin-dependent folliculogenesis. In vivo studies using conditional FOSL2 deletion in mouse granulosa cells corroborated these results, demonstrating a complete halt in GTH-dependent folliculogenesis and resultant infertility. Mechanistic exploration unveiled that FSH/FSHR initiates FOSL2 expression via the cAMP/PKA/CREB cascade, while FOSL2 in turn enhances FSHR transcription through direct promoter binding, thereby establishing a self-amplifying loop. This loop represents a molecular switch, modeled to the all-or-nothing dynamics of GTH-dependent folliculogenesis. The evolutionary conservation of this mechanism was confirmed through cross-species analyses in sheep, where FOSL2 deficiency similarly attenuated FSH/FSHR signaling and inhibited follicular growth. Our findings advance the understanding of folliculogenesis by revealing a novel FOSL2-centered amplification loop for FSH/FSHR signaling, highlighting the indispensable role of FOSL2 in reproductive biology.

cell biology↗

Kinase-Inhibitor Binding Affinity Prediction with Pretrained Graph Encoder and Language Model

MotivationThe accurate prediction of inhibitor-kinase binding affinity is crucial in drug discovery and medical applications, especially in the treatment of diseases such as cancer. Existing methods for predicting inhibitor-kinase affinity still face challenges including insufficient data expression, limited feature extraction, and low performance. Despite the progress made through artificial intelligence (AI) methods, especially deep learning technology, many current methods fail to capture the intricate interactions between kinases and inhibitors. Therefore, it is necessary to develop more advanced methods to solve the existing problems in inhibitor-kinase binding prediction. ResultsThis study proposed Kinhibit, a novel framework for inhibitor-kinase binding affinity predictor. Kinhibit integrates self-supervised pre-trained molecular encoders and protein language models (ESM-S) to extract features effectively. Kinhibit also employed a feature fusion approach to optimize the fusion of inhibitor and kinase features. Experimental results demonstrate the superiority of this method, achieving an accuracy of 92.6% in inhibitor prediction tasks of three MAPK signaling pathway kinases: Raf protein kinase (RAF), Mitogen-activated protein kinase kinase (MEK), and Extracellular Signal-Regulated Kinase (ERK). Furthermore, the framework achieves an impressive accuracy of 93.4% on a dataset containing over 200 kinases. This study provides promising and effective tools for drug screening and biological sciences.

bioinformatics↗

Follicular mural granulosa cells stockpile glycogen to fuel corpus luteum pre-vascularization

The corpus luteum (CL) arises from the luteinization of follicular granulosa cells (GCs) and theca cells, marked by rapid progesterone elevation and angiogenesis. Intriguingly, angiogenesis lags behind progesterone elevation, creating an avascular phase during which luteal cells must fuel intensive steroidogenesis without perfusion. How the avascular CL meets this energetic demand remains a mystery. Here, we reveal a novel cellular adaptive mechanism-GC energy storage (GCES)-that resolves this enigma. We demonstrate that upon luteinization initiation, GCs enter a metabolically quiescent state yet enhance glucose uptake via SLC2A1, converting the glucose into glycogen through the hCG (LH)-MAPK-RUNX1-Insulin signaling axis. Catabolism of this glycogen reserve supplies the energy required for the avascular CL, ensuring normal luteogenesis. GCES is evolutionarily conserved across species. Genetic or pharmacological disruption of GCES or glycogenolysis induces luteal insufficiency, whereas timely glucose administration enhances GCES, improving luteal function and optimize reproductive outcome in both mouse and ovine models. In human study, orally intake of glucose post-hCG significantly augments GCES and enhances progesterone production in women. These results advance luteal physiology by uncovering a universal reproductive principle with direct clinical implications.

developmental biology↗

Hantaan virus-derived peptides that stabilize HLA-E could abrogate inhibition of CD56dimNKG2A+ NK cells

NK cells could participate in the pathogenesis process of virus infectious diseases through the inhibitory receptor CD94/NKG2A interacting with HLA-E/virus-derived peptide complex. However, the effects and mechanisms of NKG2A-HLA-E axis-mediated NK cell responses in hemorrhagic fever with renal syndrome (HFRS) caused by Hantaan virus (HTNV) infection remain unclear. Single-cell RNA sequencing and flow cytometry were employed to analyze the phenotype and function of different NK cell subsets in HFRS patients. The K562/HLA-E cells binding assay was used for peptide affinity detection. The binding capacity of HLA-E/peptide-CD94/NKG2A was detected using ligand-receptor binding assay and tetramer staining. The cytotoxicity assay of NK cells against peptide-pulsed K562/HLA-E cells was conducted for functional evaluation. In this study, CD56dimCD16+NKG2A+ NK cells were the main subset in HFRS patients, showing activation and proliferation phenotypes with NKG2C-CD57- and the ability to secrete cytokines and cytotoxic mediators. Notably, none of the four identified HTNV epitopes presented by HLA-E could be recognized by CD94/NKG2A on CD56dimNKG2A+ NK cells. Furthermore, the subset of CD56dimNKG2A+ NK cells showed the enhanced cytolytic capacity against HTNV peptide pulsed K562/HLA-E cells ex vivo. Taken together, the findings demonstrate that HTNV-derived peptides presented by HLA-E could "abrogate" the inhibition of CD56dimNKG2A+ NK cells, contributing to the antiviral immune response in HFRS patients. Author SummaryHantaan virus (HTNV) is one of the main pathogens causing hemorrhagic fever with renal syndrome (HFRS) characterized by fever, hemorrhage, renal injury, and thrombocytopenia. Recently, the studies have shown that the interaction of human leukocyte antigen E (HLA-E) and natural-killer group 2, member A (NKG2A) inhibitory receptors could regulate the functions of NK cells, participating the pathogenesis process of virus infectious diseases. However, the role of NK cell response induced by HTNV infection in the pathogenesis of HFRS has not been completely determined. Here, the findings demonstrate that the elevated percentage of CD56dimNKG2A+ NK cell subset in peripheral blood of HFRS patients might exert antiviral effects through the unrecognize between CD94/NKG2A and HLA-E/HTNV peptide complex, which may abrogate the inhibition of NKG2A-expressing NK cells. This study may provide the mechanisms of NKG2A-HLA-E axis on regulating NK cell responses in HTNV infections.

immunology↗

rNMPID: a database for riboNucleoside Mono-Phosphates In DNA

MotivationRibonucleoside monophosphates (rNMPs) are the most abundant non-standard nucleotides embedded in genomic DNA. If the presence of rNMP in DNA cannot be controlled, it can lead to genome instability. The actual positive functions of rNMPs in DNA remain mainly unknown. Considering the association between rNMPs embedment and various diseases and cancer, the phenomenon of rNMPs embedment in DNA has become a prominent area of research in recent years. ResultsWe introduce the rNMPID database, which is the first database revealing rNMP-embedment characteristics, strand bias, and preferred incorporation patterns in the genomic DNA of samples from bacterial to human cells of different genetic backgrounds. The rNMPID database uses datasets generated by different rNMP-mapping techniques. It provides the researchers with a solid foundation to explore the features of rNMPs embedded in the genomic DNA of multiple sources, and their association with cellular functions, and, in future, disease. It also significantly benefits researchers in the fields of genetics and genomics who aim to integrate their studies with the rNMP-embedment data. AvailabilityrNMPID is freely accessible on the web at https://www.rnmpid.org. Contactxph6113@gmail.com or storici@gatech.edu

bioinformatics↗

Analysis of genetic signatures of tumor microenvironment yields insight into mechanisms of resistance to immunotherapy

BackgroundTherapeutic intervention targeting immune cells have led to remarkable improvements in clinical outcomes of tumor patients. However, responses are not universal. The inflamed tumor microenvironment has been reported to correlate with response in tumor patients. However, due to the lack of appropriate experimental methods, the reason why the immunotherapeutic resistance still existed on the inflamed tumor microenvironment remains unclear. Materials and methodsHere, based on integrated single-cell RNA sequencing technology, we classified tumor microenvironment into inflamed immunotherapeutic responsive and inflamed non-responsive. Then, phenotype-specific genes were identified to show mechanistic differences between distant TME phenotypes. Finally, we screened for some potential favorable TME phenotypes transformation drugs to aid current immunotherapy. ResultsMultiple signaling pathways were phenotypes-specific dysregulated. For example, Interleukin signaling pathways including IL-4 and IL-13 were activated in inflamed TME across multiple tumor types. PPAR signaling pathways and multiple epigenetic pathways were respectively inhibited and activated in inflamed immunotherapeutic non-responsive TME, suggesting a potential mechanism of immunotherapeutic resistance and target for therapy. We also identified some genetic markers of inflamed non-responsive or responsive TME, some of which have shown its potentials to enhance the efficacy of current immunotherapy. ConclusionThese results may contribute to the mechanistic understanding of immunotherapeutic resistance and guide rational therapeutic combinations of distant targeted chemotherapy agents with immunotherapy.

immunology↗

Pan-cancer analysis identified inflamed microenvironment associated multi-omics signatures

BackgroundImmunotherapy has revolutionized cancer therapy. However, responses are not universal. The inflamed tumor microenvironment has been reported to correlate with response in tumor patients. However, how different tumors shape their tumor microenvironment remains a critical unsolved problem. A deeper insight into the molecular characteristics of inflamed tumor microenvironment may be needed. Materials and methodsHere, based on single-cell RNA sequencing technology and TCGA pan-cancer cohort, we investigated multi-omics molecular features of tumor microenvironment phenotypes. Based on single-cell RNA-seq analysis, we classified pan-cancer tumor samples into inflamed or non-inflamed tumor and identified molecular features of these tumors. Analysis of integrating identified gene signatures with a drug-genomic perturbation database identified multiple drugs which may be helpful for converting non-inflamed tumors to inflamed tumors. ResultsOur results revealed several inflamed/non-inflamed tumor microenvironments-specific molecular characteristics. For example, inflamed tumors highly expressed miR-650 and lncRNA including MIR155HG and LINC00426, these tumors showed activated cytokines-related signaling pathways. Interestingly, non-inflamed tumors tended to express several genes related to neurogenesis. Multi-omics analysis demonstrated the neuro phenotype transformation may be induced by hypomethylated promoters of these genes and down-regulated miR-650. Drug discovery analysis revealed histone deacetylase inhibitors may be a potential choice for helping favorable tumor microenvironment phenotype transformation and aiding current immunotherapy. ConclusionOur results provide a comprehensive molecular-level understanding of tumor cell-immune cell interaction and may have profound clinical implications.

immunology↗