bioRxiv Science⌕ Search

Biology subjects

Xu, L.-W.

Publications and source records attributed to Xu, L.-W..

3 recordsLinked to original sources

CAPTAIN: A multimodal foundation model pretrained on co-assayed single-cell RNA and protein

Proteins act as the terminal effectors of cellular function, encoding the phenotypic consequences of genomic and transcriptomic programs. Although transcriptomic profiles serve as accessible proxies, they remain incomplete surrogates for the proteomic landscape that ultimately defines cellular phenotypes. Current single-cell foundation models, however, are trained exclusively on transcriptomes, resulting in biased and partial characterizations of cellular states. To address this limitation, we introduce CAPTAIN, a multimodal foundational model pretrained on over four million single cells with concurrently measured transcriptomes and a curated repertoire of 382 surface proteins across diverse human and mouse tissues. Our results show that CAPTAIN learns unified multimodal representations by modeling cross-modality dependencies and capturing the diversity of cellular states across complex biological contexts. CAPTAIN generalizes robustly across both fine-tuning and zero-shot settings, excelling in core downstream tasks such as protein imputation and expansion, cell type annotation, and batch harmonization. Beyond improved accuracy in multi-omics integration, CAPTAIN uncovers previously inaccessible mechanisms of protein-driven intercellular dynamics, including immune interaction patterns linked to COVID-19 severity. CAPTAIN establishes a new paradigm for multimodal single-cell modeling, laying the foundation for comprehensive cellular understanding and virtual cell construction.

bioinformatics↗

Multi-view graph learning for deciphering the dominant cell communication assembly of downstream functional events from single-cell RNA-seq data

Cell-cell communications (CCCs) from multiple sender cells collaboratively affect downstream functional events in receiver cells, thus influencing cell phenotype and function. How to rank the importance of these CCCs and find the dominant ones in a specific downstream functional event has great significance for deciphering various physiological and pathogenic processes. To date, several computational methods have been developed to focus on the identification of cell types that communicate with enriched ligand-receptor interactions from single-cell RNA-seq (scRNA-seq) data, but to the best of our knowledge, all of them lack the ability to identify the communicating cell type pairs that play a major role in a specific downstream functional event, which we call it "dominant cell communication assembly (DCA)". Here, we proposed scDCA, a multi-view graph learning method for deciphering DCA from scRNA-seq data. scDCA is based on a multi-view CCC network by constructing different cell type combinations at single-cell resolution. Multi-view graph convolution network was further employed to reconstruct the expression pattern of target genes or the functional states of receiver cells. The DCA was subsequently identified by interpreting the model with the attention mechanism. scDCA was verified in a real scRNA-seq cohort of advanced renal cell carcinoma, accurately deciphering the DCA that affect the expression patterns of the critical immune genes and functional states of malignant cells. Furthermore, scDCA also accurately explored the alteration in cell communication under clinical intervention by comparing the DCA for certain cytotoxic factors between patients with and without immunotherapy. scDCA is free available at: https://github.com/pengsl-lab/scDCA.git.

bioinformatics↗

SpaCCC: Large language model-based cell-cell communication inference for spatially resolved transcriptomic data

Drawing parallels between linguistic constructs and cellular biology, large language models (LLMs) have achieved remarkable success in diverse downstream applications for single-cell data analysis. However, to date, it still lacks methods to take advantage of LLMs to infer ligand-receptor (LR)-mediated cell-cell communications for spatially resolved transcriptomic data. Here, we propose SpaCCC to facilitate the inference of spatially resolved cell-cell communications, which relies on our fine-tuned single-cell LLM and functional gene interaction network to embed ligand and receptor genes expressed in interacting individual cells into a unified latent space. The LR pairs with a significant closer distance in latent space are taken to be more likely to interact with each other. After that, the molecular diffusion and permutation test strategies are respectively employed to calculate the communication strength and filter out communications with low specificities. The benchmarked performance of SpaCCC is evaluated on real single-cell spatial transcriptomic datasets with remarkable superiority over other methods. SpaCCC also infers known LR pairs concealed by existing aggregative methods and then identifies communication patterns for specific cell types and their signalling pathways. Furthermore, spaCCC provides various cell-cell communication visualization results at both single-cell and cell type resolution. In summary, spaCCC provides a sophisticated and practical tool allowing researchers to decipher spatially resolved cell-cell communications and related communication patterns and signalling pathways based on spatial transcriptome data.

bioinformatics↗