bioRxiv Science⌕ Search

Biology subjects

Halle, L.

Publications and source records attributed to Halle, L..

3 recordsLinked to original sources

Nicheformer: a foundation model for single-cell and spatial omics

Tissue makeup relies fundamentally on the cellular microenvironment. Spatial single-cell genomics allows probing the underlying cellular interactions in an unbiased, scalable fashion. To learn a unified cell representation that accounts for local dependencies in the cellular microenvironment, we propose Nicheformer, a transformer-based foundation model that combines human and mouse dissociated single-cell and targeted spatial transcriptomics data. Pretrained on over 57 million dissociated and 53 million spatially resolved cells across 73 tissues on cellular reconstruction, the model is fine-tuned on spatial tasks for spatial omics data to decode spatially resolved cellular information. Nicheformer excels in linear-probing and fine-tuning scenarios for a novel set of downstream tasks, in particular spatial composition prediction and spatial label prediction. We further show that existing foundation models trained on dissociated single-cell data alone are not capable of recapitulating the spatial complexity of cells in their microenvironments, indicating that multiscale models are required to understand complex local dependencies at scale. Nicheformer enables the prediction of the spatial context of dissociated cells, allowing the transfer of rich spatial information to scRNA-seq datasets. Overall, Nicheformer sets the stage for the next generation of machine-learning models in spatial single-cell analysis. Extended AbstractTissue makeup and the corresponding orchestration of vital biological activities, ranging from development and differentiation to immune response and regeneration, rely fundamentally on the cellular microenvironment and the interactions between cells. Spatial single-cell genomics allows probing such interactions in an unbiased and, increasingly, scalable fashion. To learn a unified cell representation that accounts for local dependencies in the cellular microenvironment and the underlying cell interactions, we propose to generalize recent foundation modeling approaches for disassociated single-cell transcriptomics to the spatial omics setting. Our model, Nicheformer, is a transformer-based foundation model that combines human and mouse dissociated single-cell and targeted spatial transcriptomics data to learn a cellular representation useful for a large variety of downstream tasks. Nicheformer is pretrained on over 57 million dissociated and 53 million spatially resolved cells across 73 tissues from both human and mouse. Subsequently, the model is fine-tuned on spatial tasks for spatial omics data to decode spatially resolved cellular information. We demonstrate the usefulness of Nicheformer in both linear-probing as well as fine-tuning scenarios on a novel set of spatially-relevant downstream tasks such as spatial density prediction or niche and region label prediction. In particular, we show that Nicheformer enables the prediction of the spatial context of dissociated cells, allowing the transfer of rich spatial information to scRNA-seq datasets. We define a series of novel spatial prediction problems and observe consistent top performance of Nicheformer, demonstrating the advantage of the improved model capacity of the underlying transformer. Additionally, we benchmarked Nicheformer in these tasks against scGPT1, Geneformer2, scVI3 and PCA and show that the Nicheformer architecture excels in these tasks. Altogether, our large-scale resource of more than 110 million cells in a partial spatial context, together with the set of novel spatial learning tasks and the Nicheformer model itself, will pave the way for the next generation of machine-learning models for spatial single-cell analysis.

bioinformatics↗

An integrated transcriptomic cell atlas of human endoderm-derived organoids

Human stem cells can generate complex, multicellular epithelial tissues of endodermal origin in vitro that recapitulate aspects of developing and adult human physiology. These tissues, also called organoids, can be derived from pluripotent stem cells or tissue-resident fetal and adult stem cells. However, it has remained difficult to understand the precision and accuracy of organoid cell states through comparison with primary counterparts, and to comprehensively assess the similarity and differences between organoid protocols. Advances in computational single-cell biology now allow the integration of datasets with high technical variability. Here, we integrate single-cell transcriptomes from 218 samples covering organoids of diverse endoderm-derived tissues including lung, pancreas, intestine, liver, biliary system, stomach, and prostate to establish an initial version of a human endoderm organoid cell atlas (HEOCA). The integration includes nearly one million cells across diverse conditions, data sources and protocols. We align and compare cell types and states between organoid models, and harmonize cell type annotations by mapping the atlas to primary tissue counterparts. To demonstrate utility of the atlas, we focus on intestine and lung, and clarify ontogenic cell states that can be modeled in vitro. We further provide examples of mapping novel data from new organoid protocols to expand the atlas, and showcase how integrating organoid models of disease into the HEOCA identifies altered cell proportions and states between healthy and disease conditions. The atlas makes diverse datasets centrally available, and will be valuable to assess organoid fidelity, characterize perturbed and diseased states, and streamline protocol development.

cell biology↗

hadge: a comprehensive pipeline for donor deconvolution in single cell

Single cell multiplexing techniques (cell hashing and genetic multiplexing) allow to combine multiple samples, thereby optimizing sample processing and reducing batch effects. Cell hashing conjugates antibody-tags or chemical-oligonucleotides to cell membranes, while genetic multiplexing allows to mix genetically diverse samples and relies on aggregation of RNA reads at known genomic coordinates. We developed hadge (hashing deconvolution combined with genotype information), a Nextflow pipeline that combines 12 methods to perform both hashing- and genotype-based deconvolution. We propose a joint deconvolution strategy combining the best performing methods and we demonstrate how this approach leads to recovery of previously discarded cells in a nuclei hashing of fresh-frozen brain tissue.

bioinformatics↗