bioRxiv · 10.1101/2022.09.22.509063
SIMBSIG: Similarity search and clustering for biobank-scale data
Abstract
SummaryIn many modern bioinformatics applications, such as statistical genetics, or single-cell analysis, one frequently encounters datasets which are orders of magnitude too large for conventional in-memory analysis. To tackle this challenge, we introduce SIMBSIG, a highly scalable Python package which provides a scikit-learn-like interface for out-of-core, GPU-enabled similarity searches, principal component analysis, and clustering. Due to the PyTorch backend it is highly modular and particularly tailored to many data types with a particular focus on biobank data analysis. AvailabilitySIMBSIG is freely available from PyPI and its source code and documentation can be found on GitHub (https://github.com/BorgwardtLab/simbsig) under a BSD-3 license. Contactmichael.adamer@bsse.ethz.ch
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adamer, M. F., Roellin, E., Bourguignon, L., Borgwardt, K.. 2022-09-23. SIMBSIG: Similarity search and clustering for biobank-scale data. https://doi.org/10.1101/2022.09.22.509063
Cite the original work for its findings. Save a collection to share your selection of sources.