bioRxiv · 10.1101/527515
The number of spaced-word matches between two DNA sequences as a function of the underlying pattern weight
Abstract
We study the number Nk of (spaced) word matches between pairs of evolutionarily related DNA sequences depending on the word length or pattern weight k, respectively. We show that, under the Jukes-Cantor model, the number of substitutions per site that occurred since two sequences evolved from their last common ancestor, can be esti-mated from the slope of a certain function of Nk. Based on these considerations, we implemented a software program for alignment-free sequence comparison called Slope-SpaM. Test runs on simulated sequence data show that Slope-SpaM can estimate phylogenetic dis-tances with high accuracy for up to around 0.5 substitutions per po-sitions. The statistical stability of our results is improved if spaced words are used instead of contiguous k-mers. Unlike previous methods that are based on the number of (spaced) word matches, our approach can deal with sequences that share only local homologies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Röhling, S., Morgenstern, B.. 2019-01-23. The number of spaced-word matches between two DNA sequences as a function of the underlying pattern weight. https://doi.org/10.1101/527515
Cite the original work for its findings. Save a collection to share your selection of sources.