bioRxiv · 10.1101/2020.10.01.322164
MetaGraph: Indexing and Analysing Nucleotide Archives at Petabase-scale
Abstract
The amount of biological sequencing data available in public repositories is growing exponentially, forming an invaluable biomedical research resource. Yet, making it full-text searchable and easily accessible to researchers in life and data science is an unsolved problem. In this work, we take advantage of recently developed, very efficient data structures and algorithms for representing sequence sets. We make Petabases of DNA sequences across all clades of life, including viruses, bacteria, fungi, plants, animals, and humans, fully searchable. Our indexes are freely available to the research community. This highly compressed representation of the input sequences (up to 5800x) fits on a single consumer hard drive ({approx}100 USD), making this valuable resource cost-effective to use and easily transportable. We present the underlying methodological framework, called MetaGraph, that allows us to scalably index very large sets of DNA or protein sequences using annotated De Bruijn graphs. We demonstrate the feasibility of indexing the full extent of existing sequencing data and present new approaches for efficient and cost-effective full-text search at an on-demand cost of $0.10 per queried Mpb. We explore several practical use cases to mine existing archives for interesting associations and demonstrate the utility of our indexes for integrative analyses.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Karasikov, M., Mustafa, H., Danciu, D., Zimmermann, M., Barber, C., Ratsch, G., Kahles, A.. 2020-10-02. MetaGraph: Indexing and Analysing Nucleotide Archives at Petabase-scale. https://doi.org/10.1101/2020.10.01.322164
Cite the original work for its findings. Save a collection to share your selection of sources.