bioRxiv · 10.1101/2022.01.15.476407
scSampler: fast diversity-preserving subsampling of large-scale single-cell transcriptomic data
Abstract
SummaryThe number of cells measured in single-cell transcriptomic data has grown fast in recent years. For such large-scale data, subsampling is a powerful and often necessary tool for exploratory data analysis. However, the easiest random subsampling is not ideal from the perspective of preserving rare cell types. Therefore, diversity-preserving subsampling is required for fast exploration of cell types in a large-scale dataset. Here we propose scSampler, an algorithm for fast diversity-preserving subsampling of single-cell transcriptomic data. Using simulated and real data, we show that scSampler consistently outperforms existing subsam-pling methods in terms of both the computational time and the Hausdorff distance between the full and subsampled datasets. AvailabilityscSampler is implemented in Python and is published under the MIT source license. It can be installed by pip install scsampler and used with the Scanpy pipline. The code is available on GitHub: https://github.com/SONGDONGYUAN1994/scsampler. Contactlinwang@gwu.edu; jli@stat.ucla.edu
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Song, D., Xi, N. M., Li, J. J., Wang, L.. 2022-01-18. scSampler: fast diversity-preserving subsampling of large-scale single-cell transcriptomic data. https://doi.org/10.1101/2022.01.15.476407
Cite the original work for its findings. Save a collection to share your selection of sources.