bioRxiv · 10.1101/620666
Deconvolute individual genomes from metagenome sequences through short read clustering
Abstract
MotivationMetagenome assembly from short next-generation sequencing data is a challenging process due to its large scale and computational complexity. Clustering short reads before assembly offers a unique opportunity for parallel downstream assembly of genomes with individualized optimization. However, current read clustering methods suffer either false negative (under-clustering) or false positive (over-clustering) problems.\n\nResultsBased on a previously developed scalable read clustering method on Apache Spark, SpaRC, that has very low false positives, here we extended its capability by adding a new method to further cluster small clusters. This method exploits statistics derived from multiple samples in a dataset to reduce the under-clustering problem. Using a synthetic dataset from mouse gut microbiomes we show that this method has the potential to cluster almost all of the reads from genomes with sufficient sequencing coverage. We also explored several clustering parameters that deferentially affect genomes with various sequencing coverage.\n\nAvailabilityhttps://bitbucket.org/berkeleylab/jgi-sparc/.\n\nContactzhongwang@lbl.gov
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Li, K., Wang, L., Shi, L., Deng, L., Wang, Z.. 2019-04-29. Deconvolute individual genomes from metagenome sequences through short read clustering. https://doi.org/10.1101/620666
Cite the original work for its findings. Save a collection to share your selection of sources.