Systematic evaluation of robustness to cell type mismatch of deconvolution methods for spatial transcriptomics data
Sequencing-based spatial transcriptomics (ST) approaches preserve spatial information but with limited cellular resolution, whereas single-cell RNA-sequencing (scRNA-seq) techniques provide single-cell resolution but lose spatial context during tissue dissociation. Given these complementary strengths, computational tools have been developed to combine scRNA-seq and ST data. These methods use deconvolution techniques to identify cell types and estimate their proportions at each spatial location in ST data, using scRNA-seq reference data. However, these methods are sensitive to missing cell types in the scRNA-seq reference, a problem known as cell type mismatch. Using two reference datasets, we performed extensive simulations to systematically evaluate the robustness to cell type mismatch of six deconvolution methods (CARD, cell2location, RCTD, Seurat, SPOTlight, Stereoscope) tailored for ST data, and two designed for bulk RNA-seq data (MuSiC, SCDC). At baseline, that is, with no cell types missing from the reference datasets, cell2location showed the strongest performance, while Seurat performed the worst. By simulating different cell type mismatch scenarios, we found that the performance of deconvolution methods decreases proportionally to the number of cell types missing from the reference. Moreover, compared to baseline, for most methods the relative decrease in performance is similar. Additionally, methods that perform well at baseline tend to assign the proportions of a missing cell type to the transcriptionally most similar cell types present in the reference data. Our results highlight the adverse effects of cell type mismatch on the performance of deconvolution methods for ST data and stress the need for more robust approaches to this issue.