bioRxiv · 10.64898/2026.04.11.717925
GraphMana: graph-native data management for population genomics projects
Abstract
Population genomics projects rely on fragmented file-based workflows that lose provenance and require full reprocessing when samples are added. Graph-Mana stores variant data in a graph database as packed genotype arrays with pre-computed population statistics, enabling incremental sample addition, provenance tracking, cohort management, and export to 17 formats. Two access paths serve different needs: a FAST PATH reading population-level arrays in O(K) time and a FULL PATH unpacking per-sample genotypes in O(N) time. On human 1000 Genomes data (3,202 samples, 70.7M variants), Graph-Mana completed a 46-operation lifecycle in 98 minutes from a single persistent database.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Estaji, E., Zhao, S.-W., Chen, Z.-Y., Nie, S., Mao, J.-F.. 2026-04-14. GraphMana: graph-native data management for population genomics projects. https://doi.org/10.64898/2026.04.11.717925
Cite the original work for its findings. Save a collection to share your selection of sources.