bioRxiv · 10.1101/2025.08.23.671699
Adding layers of information to scRNA-seq data using pre-trained language models
Abstract
Pre-trained language models promise to enrich single-cell analyses with contextual information from large biomedical text corpora, but it remains unclear how to optimally align this knowledge with quantitative scRNA-seq data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then fine-tune lightweight encoder-only biomedical language models to learn a shared, literature-enriched representation. Controlled evaluations across immune and developmental datasets show that this representation preserves cell identity while adding robust and interpretable contextual layers of functional, disease-associated, and developmental information to single-cell analysis workflows.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Krissmer, S. M., Menger, J., Rollin, J., Vogel, T. M., Binder, H., Hackenberg, M.. 2025-08-27. Adding layers of information to scRNA-seq data using pre-trained language models. https://doi.org/10.1101/2025.08.23.671699
Cite the original work for its findings. Save a collection to share your selection of sources.