bioRxiv · 10.1101/588582
Magnus Representation of Genome Sequences
Abstract
We introduce an alignment-free method, the Magnus Representation, to analyze genome sequences. The Magnus Representation captures higher-order information in genome sequences. We combine our approach with the idea of k-mers to define an effectively computable Mean Magnus Vector. We perform phylogenetic analysis on three datasets: mosquito-borne viruses, filoviruses, and bacterial genomes. Our results on ebolaviruses are consistent with previous phylogenetic analyses, and confirm the modern viewpoint that the 2014 West African Ebola outbreak likely originated from Central Africa. Our analysis also confirms the close relationship between Bundibugyo ebolavirus and Tai Forest ebolavirus. For bacterial genomes, our method is able to classify relatively well at the family and genus level, as well as at higher levels such as phylum level. The bacterial genomes are also separated well into Gram-positive and Gram-negative subgroups.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wu, C., Ren, S., Wu, J., Xia, K.. 2019-03-25. Magnus Representation of Genome Sequences. https://doi.org/10.1101/588582
Cite the original work for its findings. Save a collection to share your selection of sources.