bioRxiv · 10.1101/2023.07.06.548007
Topological stratification of continuous genetic variation in large biobanks
Abstract
Biobanks now contain genetic data from millions of individuals. Dimensionality reduction, visualization and clustering are standard when exploring data at these scales; while efficient and tractable methods exist for the first two, clustering remains challenging because of uncertainty about sources of population structure. In practice, clustering is commonly performed by drawing shapes around dimensionally reduced data or assuming populations have a "type" genome. We propose a method of clustering data with topological analysis that is fast, easy to implement, and integrates with existing pipelines. The approach is robust to the presence of sub-populations of varying sizes and wide ranges of population structure patterns. We use UMAP and HDBSCAN, respectively methods of dimensionality reduction and density clustering, on data from three biobanks. We illustrate how topological genetic strata can help us understand structure within biobanks, evaluate distributions of genotypic and phenotypic data, examine polygenic score transferability, identify potential influential alleles, and perform quality control.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Diaz-Papkovich, A., Zabad, S., Ben-Eghan, C., Anderson-Trocme, L., Femerling, G., Nathan, V., Patel, J., Gravel, S.. 2023-07-07. Topological stratification of continuous genetic variation in large biobanks. https://doi.org/10.1101/2023.07.06.548007
Cite the original work for its findings. Save a collection to share your selection of sources.