bioRxiv Science⌕ Search

Biology subjects

Cavinato, T.

Publications and source records attributed to Cavinato, T..

3 recordsLinked to original sources

Evaluating anonymized genome re-identification using polygenic predictions and its implications for data privacy

Re-identification by phenotypic prediction aims to determine whether a genome belongs to a specific individual by comparing the individuals known traits with those predicted from the genome. This type of tracing attack is widely discussed in the genomic privacy literature, yet previous studies have been criticized for overstating its practical risks. Over the past decade, genome-wide association studies (GWAS) with increasing sample size improved the accuracy of phenotypic prediction, potentially enhancing such attacks. To quantify their real-world threat, we developed a probabilistic framework that estimates the likelihood of a match between an individuals observed traits and polygenic scores (PGS) derived from a genome, while accounting for prediction accuracy and genetic and environmental correlations between the traits. We benchmarked this re-identification method and examined how the prior probability (reflecting the a priori chance that a random genome and set of traits correspond to the same person) affects performance. Finally, we assessed whether sensitive information could be inferred through this attack by attempting to predict multiple sensitive haplotypes, such as APOE-{varepsilon}4 (linked with Alzheimers disease). Our re-identification method outperformed a state-of-the-art tool, and reached a precision above 99% for a recall of 40% when considering a prior of 50%. However, after considering real-world settings, we estimated that realistic priors would not exceed 4 x 10-4%, resulting in a precision lower than 0.13% at the same recall (40%). The inference of sensitive genotypes also proved ineffective, as achieving a precision above 50% for identifying APOE-{varepsilon}4 carriers was only possible at a recall below 20%. To conclude, although re-identification by phenotypic prediction is technically feasible, our findings indicate that its effectiveness in real-world conditions is limited. These results counterpoint to earlier claims of severe genomic privacy risks and offer guidance for policymakers, biobank administrators, and research participants.

genetics↗

Parental haplotypes reconstruction in up to 440,209 individuals reveals recent assortative mating dynamics

Assortative mating (AM), the tendency for individuals to choose partners with similar traits, plays an important role in social stratification and has a wide-spread impact on the genetic architecture of complex human traits. A genetic footprint of this behaviour, genetic assortative mating (GAM), has been documented for many traits. However, most existing approaches to estimate GAM either rely on genotyped couples - rare in large biobanks - or on estimates based on gametic phase disequilibrium (GPD), which reflect cumulative effects over multiple generations and cannot detect short-term events. We introduce a novel, scalable approach that infers genetic data for the parental generation to boost statistical power and improve interpretability when estimating GAM in biobank cohorts. Our method reconstructs the two parental haplotypes of biobank individuals by leveraging state-of-the-art inter-chromosomal phasing based on close relative information. By correlating polygenic scores computed separately on maternally and paternally inherited haplotypes, we infer the extent of GAM. Applied to 245,884 individuals from the UK Biobank and 194,325 individuals from the Estonian Biobank, our haplotype-based estimates showed strong concordance with GAM estimates from genotyped mate pairs across 69 traits and improved performance compared to existing GPD-based methods. We replicated previously identified patterns of GAM for many traits (educational attainment, height, BMI, and alcohol consumption), and revealed new ones (overall health and sedentary lifestyle). Temporal and geographic stratification revealed accelerated assortment in recent generations -- particularly for education and height -- and modest differences between urban and rural contexts. Our method enables scalable, interpretable, and generation-specific estimation of GAM in large biobank cohorts, providing new insights into the dynamic nature of mate choice and its impact on the genetic makeup of populations.

genetics↗

A resampling-based approach to share reference panels

For many genome-wide association studies, imputing genotypes from a haplotype reference panel is a necessary step. Over the past 15 years, reference panels have become larger and more diverse, leading to improvements in imputation accuracy. However, the latest generation of reference panels is subject to restrictions on data sharing due to concerns about privacy, limiting their usefulness for genotype imputation. In this context, we propose RESHAPE, a method that employs a recombination Poisson process on a reference panel to simulate the genomes of hypothetical descendants after multiple generations. This data transformation helps to protect against re-identification threats and preserves important data attributes, such as linkage disequilibrium (LD) patterns and, to some degree, Identity-By-Descent (IBD) sharing, allowing for genotype imputation. Our experiments on gold standard datasets show that simulated descendants up to eight generations can serve as reference panels without significantly reducing genotype imputation accuracy. We suggest that this specific type of data anonymization could be used to generate synthetic reference panels available under less restrictive data sharing policies.

bioinformatics↗