bioRxiv · 10.1101/2022.12.19.521123
Genomic Distance-based Rapid Uncovering of Microbial Population Structures (GRUMPS): a reference free genomic data cleaning methodology
Abstract
Accurate datasets are crucial for rigorous large-scale sequence-based analyses such as those performed in phylogenomics and pangenomics. As the volume of available sequence data grows and the quality of these sequences varies, there is a pressing need for reliable methods to swiftly identify and eliminate low-quality and misidentified genomes from datasets prior to analysis. Here we introduce a robust, controlled, computationally efficient method for deriving species-level population structures of bacterial species, regardless of the dataset size. Additionally, our pipeline can classify genomes into their respective species at the genus level. By leveraging this methodology, researchers can rapidly clean datasets encompassing entire bacterial species and examine the sub-species population structures within the provided genomes. These cleaned datasets can subsequently undergo further refinement using a variety of methods to yield sequence sets with varying levels of diversity that faithfully represent entire species. Increasing the efficiency and accuracy of curation of species-level datasets not only enhances the reliability of downstream analyses, but also facilitates a deeper understanding of bacterial population dynamics and evolution.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abram, K. Z., Udaondo, Z., Nookaew, I., Robeson, M. S., Jun, S.-R.. 2022-12-20. Genomic Distance-based Rapid Uncovering of Microbial Population Structures (GRUMPS): a reference free genomic data cleaning methodology. https://doi.org/10.1101/2022.12.19.521123
Cite the original work for its findings. Save a collection to share your selection of sources.