bioRxiv · 10.1101/729665
Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data
Abstract
We proposed a new sequence sketching technique named k-mer substring space decomposition (kssd), which sketches sequences via k-mer substring space sampling instead of local-sensitive hashing. Kssd is more accurate and faster for resemblance estimation than other sketching methods developed so far. Notably, kssd is robust even when two sequences are of very different sizes. For containment analysis, kssd slightly outperformed mash screen--its closest competitor--in accuracy, while took testing datasets of 110,535 times less space occupation and consumed 2,523 times less CPU time than mash screen--suggesting kssd is suite for quick containment analysis for almost the entire omics datasets deposited in NCBI. We detailed the kssd algorithm, provided proofs of its statistical properties and discussed the roots of its superiority, limitations and future directions. Kssd is freely available under an Apache License, Version 2.0 (https://github.com/yhg926/public_kssd)
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yi, H., Lin, Y., Jin, W.. 2019-08-14. Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data. https://doi.org/10.1101/729665
Cite the original work for its findings. Save a collection to share your selection of sources.