bioRxiv · 10.1101/223404
Computational haplotype recovery and long-read validation identifies novel isoforms of industrially relevant enzymes from natural microbial communities
Abstract
Elucidation of population-level diversity of microbiomes is a significant step towards a complete understanding of the evolutionary, ecological and functional importance of microbial communities. Characterizing this diversity requires the recovery of the exact DNA sequence (haplotype) of each gene isoform from every individual present in the community. To address this, we present Hansel and Gretel: a freely-available data structure and algorithm, providing a software package that reconstructs the most likely haplotypes from metagenomes. We demonstrate recovery of haplotypes from short-read Illumina data for a bovine rumen microbiome, and verify our predictions are 100% accurate with long-read PacBio CCS sequencing. We show that Gretels haplotypes can be analyzed to determine a significant difference in mutation rates between core and accessory gene families in an ovine rumen microbiome. All tools, documentation and data for evaluation are open source and available via our repository: https://github.com/samstudio8/gretel
Explore related subjects
Keep this discovery
Nicholls, S. M., Aubrey, W., Edwards, A., de Grave, K., Huws, S., Schietgat, L., Soares, A., Creevey, C. J., Clare, A.. 2017-11-22. Computational haplotype recovery and long-read validation identifies novel isoforms of industrially relevant enzymes from natural microbial communities. https://doi.org/10.1101/223404
Cite the original work for its findings. Save a collection to share your selection of sources.