bioRxiv · 10.1101/2021.12.16.473007
anndata: Annotated data
Abstract
anndata is a Python package for handling annotated data matrices in memory and on disk (github.com/theislab/anndata), positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and a PyTorch interface. Statement of needGenerating insight from high-dimensional data matrices typically works through training models that annotate observations and variables via low-dimensional representations. In exploratory data analysis, this involves iterative training and analysis using original and learned annotations and task-associated representations. anndata offers a canonical data structure for book-keeping these, which is neither addressed by pandas (McKinney, 2010), nor xarray (Hoyer & Hamman, 2017), nor commonly-used modeling packages like scikit-learn (Pedregosa et al., 2011).
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Virshup, I., Rybakov, S., Theis, F. J., Angerer, P., Wolf, F. A.. 2021-12-19. anndata: Annotated data. https://doi.org/10.1101/2021.12.16.473007
Cite the original work for its findings. Save a collection to share your selection of sources.