bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.10.29.620828

MARVEL: Microenvironment Annotation by Supervised Graph Contrastive Learning

Abstract

Recent advancements in in situ molecular profiling technologies, including spatial proteomics and transcriptomics, have enabled detailed characterization of the microenvironment at cellular and subcellular levels. While these techniques provide rich information about individual cells spatial coordinates and expression profiles, extracting biologically meaningful spatial structures from the data remains a significant challenge. Current methodologies often rely on unsupervised clustering followed by cell type annotation based on differentially expressed genes within each cluster and most of the time will require other information as the reference (e.g., HE-stained images). This is labor-intensive and demands extensive domain knowledge. To address these challenges, we propose a supervised graph contrastive learning framework, MARVEL. MARVEL is a supervised graph contrastive learning method that can effectively embed local microenvironments represented by cell neighbor graphs into a continuous representation space, facilitating various downstream microenvironment annotation scenarios. By leveraging partially annotated examples as strong positives, our approach mitigates the common issues of false positives encountered in conventional graph contrastive learning. Using real-world annotated data, we demonstrate that MARVEL outperforms existing methods in three key microenvironment-related tasks: transductive microenvironment annotation, inductive microenvironment querying, and the identification of novel microenvironments across different slices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

CUI, Y., Wen, H., Yang, R., Luo, X., Liu, H., Xie, Y.. 2024-11-03. MARVEL: Microenvironment Annotation by Supervised Graph Contrastive Learning. https://doi.org/10.1101/2024.10.29.620828

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Mapping Gene Expression to an Interpretable Semantic Space

Cell embeddings organize single-cell expression data, but their dimensions have no biological meaning, so clusters are interpreted afterward. We present MESIC (Mapping Expression to Semantic space with Interpretable Components), which builds the written knowledge about genes held in curated databases into the dimensions themselves. A biomedical language model converts each gene's summary into a semantic embedding. MESIC compresses these embeddings into a small number of components, each concentrated on a small set of genes and explained by their annotations. The components are computed once from the summaries, so any expression dataset can be mapped onto them, and every cluster, outlier, or cell-type assignment is then characterized by named genes. In cardiomyocytes, outliers in the component space were enriched for hypertrophic cardiomyopathy. In a lung atlas, unsupervised clusters in that space matched the broad cell types that experts had annotated. In both, the components that separated the cells matched their known biology. For about half of the cells that the atlas itself had left unannotated, the same space gave a confident cluster assignment, and with it an interpretation through component-associated genes. Gene summaries thus give single-cell analysis a coordinate system in which every result is traced to genes and what is written about them.

bioinformatics↗

Physical priors improve performance of structure-based binding affinity models

Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.

bioinformatics↗

Rapid Shift Toward Pulsed Field Ablation and Precision Risk Stratification in High-Impact Atrial Fibrillation Research

Conventional bibliometrics rely on lifetime citations, obscuring immediate shifts in cardiovascular research paradigms. To track emerging trends in atrial fibrillation management, we performed a comparative bibliometric analysis of the 50 highest-cited original research articles per year in OpenAlex topic T10065 across consecutive 2023 (Class of 2025) and 2024 (Class of 2026) publication cohorts. Articles were ranked using a fixed 18-month post-publication citation window, and extracted concepts were normalized into 719 canonical topics and 32 parent themes using large language model curation. Concept frequency tracking demonstrated a swift technological shift, with pulsed field ablation showing the largest topic frequency increase (+0.08, from 0.30 to 0.38) to become the leading canonical topic in 2024, displacing conventional thermal pulmonary vein isolation (-0.18, 0.40 to 0.22). Simultaneously, stroke prevention focus shifted toward refined predictive modeling, with increases in Risk Stratification and Predictive Models (+0.08, 0.30 to 0.38) and CHA2DS2-VASc scoring (+0.08, from 0.10 to 0.18). High-impact atrial fibrillation research is rapidly pivoting toward non-thermal ablation safety profiling and precision risk stratification, highlighting the utility of fixed-window concept mining for capturing real-time scientific evolution. Online explorer of the result is available at https://pri.pepkio.com.

bioinformatics↗