bioRxiv · 10.1101/2024.11.28.625845
On metrics for subpopulation detection in single-cell and spatial omics data
Abstract
Benchmarks are crucial to understanding the strengths and weaknesses of the growing number of tools for single-cell and spatial omics analysis. A key task is to distinguish subpopulations within complex tissues, where evaluation typically relies on external clustering validation metrics. Different metrics often lead to inconsistencies between rankings, highlighting the importance of understanding the behavior and biological implications of each metric. In this work, we provide a framework for systematically understanding and selecting validation metrics for single-cell data analysis, addressing tasks such as creating cell embeddings, constructing graphs, clustering, and spatial domain detection. Our discussion centers on the desirable properties of metrics, focusing on biological relevance and potential biases. Using this framework, we not only analyze existing metrics, but also develop novel ones. Delving into domain detection in spatial omics data, we develop new external metrics tailored to spatially-aware measurements. Additionally, a Bioconductor R package, poem, implements all the metrics discussed. While we focus on single-cell omics, much of the discussion is of broader relevance to other types of high-dimensional data.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Luo, S., Germain, P.-L., von Meyenn, F., Robinson, M. D.. 2024-12-03. On metrics for subpopulation detection in single-cell and spatial omics data. https://doi.org/10.1101/2024.11.28.625845
Cite the original work for its findings. Save a collection to share your selection of sources.