bioRxiv · 10.1101/2025.10.17.682995
PRISM-G: an interpretable privacy scoring method for assessing risk in synthetic human genome data
Abstract
Synthetic genomic data promises broader access, but unresolved privacy risks persist. In Europe, these risks increasingly hinder cross-border use of national genomic resources due to limitations in trust and legal interoperability rather than scientific demand. At the same time, privacy risk is not uniformly distributed: leakage driven by relatedness structure and rare-variant uniqueness can disproportionately affect underrepresented or vulnerable populations when shared data are misused or linked with external resources, making transparent, domain-aware measurement of privacy exposure central to responsible governance. We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genome data cross three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure via rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT-solver (Genomator). Our results show that privacy vulnerabilities concentrate along different axes across models and marker densities, demonstrating that a single similarity-based metric is sufficient to characterize genomic privacy risk.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Correa Rojo, A., Moreau, Y., Ertaylan, G.. 2025-10-17. PRISM-G: an interpretable privacy scoring method for assessing risk in synthetic human genome data. https://doi.org/10.1101/2025.10.17.682995
Cite the original work for its findings. Save a collection to share your selection of sources.