bioRxiv · 10.1101/2024.04.24.590989
SECRET-GWAS: Confidential Computing for Population-Scale GWAS
Abstract
Genomic data from a single institution lacks global diversity representation, especially for rare variants and diseases. Confidential computing can enable collaborative GWAS without compromising privacy or accuracy, however, due to limited secure memory space and performance overheads previous solutions fail to support widely used regression methods. We present SECRET-GWAS: a rapid, privacy-preserving, population-scale, collaborative GWAS tool. We discuss several system optimizations, including streaming, batching, data parallelization, and reducing trusted hardware overheads to efficiently scale linear and logistic regression to over a thousand processor cores on an Intel SGX-based cloud platform. In addition, we protect SECRET-GWAS against several hardware side-channel attacks, including Spectre, using data-oblivious code transformations and optimized speculative load hardening. SECRET-GWAS is an open-source tool and works with the widely used Hail genomic analysis framework. Our experiments on Azures Confidential Computing platform demonstrate that SECRET-GWAS enables multivariate linear and logistic regression GWAS queries on population-scale datasets (one million patients, four million SNPs, 12 covariates) from ten independent sources in just 4.5 and 29 minutes, respectively.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rosenblum, J., Dong, J., Narayanasamy, S.. 2024-04-28. SECRET-GWAS: Confidential Computing for Population-Scale GWAS. https://doi.org/10.1101/2024.04.24.590989
Cite the original work for its findings. Save a collection to share your selection of sources.