bioRxiv · 10.1101/559807
FQSqueezer: k-mer-based compression of sequencing data
Abstract
MotivationThe amount of genomic data that needs to be stored is huge. Therefore it is not surprising that a lot of work has been done in the field of specialized data compression of FASTQ files. The existing algorithms are, however, still imperfect and the best tools produce quite large archives. ResultsWe present FQSqueezer, a novel compression algorithm for sequencing data able to process single- and paired-end reads of variable lengths. It is based on the ideas from the famous prediction by partial matching and dynamic Markov coder algorithms known from the general-purpose-compressors world. The compression ratios are often tens of percent better than offered by the state-of-the-art tools. Availability and Implementationhttps://github.com/refresh-bio/fqsqueezer Contactsebastian.deorowicz@polsl.pl Supplementary informationSupplementary data are available at publishers Web site.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Deorowicz, S.. 2019-02-24. FQSqueezer: k-mer-based compression of sequencing data. https://doi.org/10.1101/559807
Cite the original work for its findings. Save a collection to share your selection of sources.