bioRxiv · 10.1101/2022.06.03.494761
Efficient Pangenome Construction through Alignment-Free Residue Pangenome Analysis (ARPA)
Abstract
Protein sequences can be transformed into vectors composed of counts for each amino acid (vector of Residue Counts; vRC) that are mathematically tractable and retain information about homology. We use vRCs to perform alignment-free, residue-based, pangenome analysis (ARPA; https://github.com/Arnavlal/ARPA). ARPA is 70-90 times faster at identifying homologous gene clusters compared to standard techniques, and offers rapid calculation, visualization, and novel phylogenetic approaches for pangenomes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lal, A., Moustafa, A. M., Planet, P.. 2022-06-05. Efficient Pangenome Construction through Alignment-Free Residue Pangenome Analysis (ARPA). https://doi.org/10.1101/2022.06.03.494761
Cite the original work for its findings. Save a collection to share your selection of sources.