bioRxiv Science⌕ Search

Biology subjects

Krissmer, S. M.

Publications and source records attributed to Krissmer, S. M..

3 recordsLinked to original sources

Coding agents author interpretable single-cell embedding models from the literature

The single-cell literature catalogs cell states as validated marker-gene programs -- a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction. Here we show coding agents can author single-cell embedding models directly from the literature. Given a scenario that focuses this literature lens on a chosen biological subdomain, the agent edits a structured Python template, curating named, literature-cited gene programs and composing them into axes, without a gene-set database, training, or sight of the data. Across mouse and human tissues these zero-shot embeddings are competitive in biological quality with conventional, foundation-model, and program-informed baselines, batch-robust by construction and reproducible across runs, complementing data-driven embeddings. Because each dimension is a named, cited gene program, the embedding is interpretable and auditable, and its composable axes can be steered into a developmental tree.

bioinformatics↗

mmContext: an open framework for multimodal contrastive learning of omics and text data

SummaryMultimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData .obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. Availability and implementationPretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training.

bioinformatics↗

Adding layers of information to scRNA-seq data using pre-trained language models

Pre-trained language models promise to enrich single-cell analyses with contextual information from large biomedical text corpora, but it remains unclear how to optimally align this knowledge with quantitative scRNA-seq data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then fine-tune lightweight encoder-only biomedical language models to learn a shared, literature-enriched representation. Controlled evaluations across immune and developmental datasets show that this representation preserves cell identity while adding robust and interpretable contextual layers of functional, disease-associated, and developmental information to single-cell analysis workflows.

bioinformatics↗