bioRxiv · 10.1101/2020.12.10.420448
Assembling Long Accurate Reads Using de Bruijn Graphs
Abstract
Although most existing genome assemblers are based on the de Bruijn graphs, it remains unclear how to construct these graphs for large genomes and large k-mer sizes. This algorithmic challenge has become particularly important with the emergence of long high-fidelity (HiFi) reads that were recently utilized to generate a semi-manual telomere-to-telomere assembly of the human genome and to get a glimpse into biomedically important regions that evaded all previous attempts to sequence them. To enable automated assemblies of long and accurate reads, we developed a fast LJA algorithm that reduces the error rate in these reads by three orders of magnitude (making them nearly error-free) and constructs the de Bruijn graph for large genomes and large k-mer sizes. Since the de Bruijn graph constructed for a fixed k-mer size is typically either too tangled or too fragmented, LJA uses a new concept of a multiplex de Bruijn graph with varying k-mer sizes. We demonstrate that LJA improves on the state-of-the-art assemblers with respect to both accuracy and contiguity and enables automated telomere-to-telomere assemblies of entire human chromosomes.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bankevich, A., Bzikadze, A. V., Kolmogorov, M., Pevzner, P.. 2020-12-11. Assembling Long Accurate Reads Using de Bruijn Graphs. https://doi.org/10.1101/2020.12.10.420448
Cite the original work for its findings. Save a collection to share your selection of sources.