bioRxiv ScienceSearch

Biology subjects

Damrauer, S. M.

Publications and source records attributed to Damrauer, S. M..

2 recordsLinked to original sources

Exome-by-phenome-wide rare variant gene burden association with electronic health record phenotypes

BackgroundBy coupling large-scale DNA sequencing with electronic health records (EHR), \"genome-first\" approaches can enhance our understanding of the contribution of rare genetic variants to disease. Aggregating rare, loss-of-function variants in a candidate gene into a \"gene burden\" to test for association with EHR phenotypes can identify both known and novel clinical implications for the gene in human disease. However, this methodology has not yet been applied on both an exome-wide and phenome-wide scale, and the clinical ontologies of rare loss-of-function variants in many genes have yet to be described.\n\nMethodsWe leveraged whole exome sequencing (WES) data in participants (N=11,451) in the Penn Medicine Biobank (PMBB) to address on an exome-wide scale the association of a burden of rare loss-of-function variants in each gene with diverse EHR phenotypes using a phenome-wide association study (PheWAS) approach. For discovery, we collapsed rare (minor allele frequency (MAF) [&le;] 0.1%) predicted loss-of-function (pLOF) variants (i.e. frameshift insertions/deletions, gain/loss of stop codon, or splice site disruption) per gene to perform a gene burden PheWAS. Subsequent evaluation of the significant gene burden associations was done by collapsing rare (MAF [&le;] 0.1%) missense variants with Rare Exonic Variant Ensemble Learner (REVEL) scores [&ge;] 0.5 into corresponding yet distinct gene burdens, as well as interrogation of individual low-frequency to common (MAF > 0.1%) pLOF variants and missense variants with REVEL[&ge;] 0.5. We replicated our findings using the UK Biobanks (UKBB) whole exome sequence dataset (N=49,960).\n\nResultsFrom the pLOF-based discovery phase, we identified 106 gene burdens with phenotype associations at p<10-6 from our exome-by-phenome-wide association studies. Positive-control associations included TTN (cardiomyopathy, p=7.83E-13), MYBPC3 (hypertrophic cardiomyopathy, p=3.48E-15), CFTR (cystic fibrosis, p=1.05E-15), CYP2D6 (adverse effects due to opiates/narcotics, p=1.50E-09), and BRCA2 (breast cancer, p=1.36E-07). Of the 106 genes, 12 gene-phenotype relationships were also detected by REVEL-informed missense-based gene burdens and 19 by single-variant analyses, demonstrating the robustness of these gene-phenotype relationships. Three genes showed evidence of association using both additional methods (BRCA1, CFTR, TGM6), leading to a total of 28 robust gene-phenotype associations within PMBB. Furthermore, replication studies in UKBB validated 30 of 106 gene burden associations, of which 12 demonstrated robustness in PMBB.\n\nConclusionOur study presents 12 exome-by-phenome-wide robust gene-phenotype associations, which include three proof-of-concept associations and nine novel findings. We show the value of aggregating rare pLOF variants into gene burdens on an exome-wide scale for unbiased association with EHR phenotypes to identify novel clinical ontologies of human genes. Furthermore, we show the significance of evaluating gene burden associations through complementary, yet non-overlapping genetic association studies from the same dataset. Our results suggest that this approach applied to even larger cohorts of individuals with WES or whole-genome sequencing data linked to EHR phenotype data will yield many new insights into the relationship of genetic variation and disease phenotypes.

genomics

Assessing a causal relationship between circulating lipids and breast cancer risk via Mendelian randomization

ObjectiveTo assess a potential causal relationship between genetic variants associated with plasma lipid traits (high-density lipoprotein cholesterol, HDL; low-density lipoprotein cholesterol, LDL; triglycerides, TG) with risk for breast cancer.\n\nDesignMendelian randomization (MR) study.\n\nSetting and ParticipantsData from genome-wide association studies in up to 215,551 subjects from the Million Veterans Project were used to construct genetic instruments for plasma lipid traits. The effect of these instruments on breast cancer risk was evaluated using genetic data from the BCAC consortium based on 122,977 breast cancer cases and 105,974 controls.\n\nExposuresGenetically modified plasma levels of LDL, HDL, or TG.\n\nMain Outcomes and MeasuresOdds ratio (OR) for breast cancer risk per standard-deviation increase in HDL, LDL, or TG.\n\nResultsWe observed that a 1-SD genetically determined increase in HDL levels is associated with an increased risk for all breast cancers (HDL: OR=1.08, 95% CI=1.04-1.13, P=7.4x10-5).\n\nMultivariable MR analysis, which adjusted for the effects of LDL, TG, body mass index, and age at menarche, corroborated this observation for HDL (OR=1.06, 95% CI=1.03-1.10, P=4.9x10-4) and also a relationship between LDL and breast cancer risk (OR=1.03, 95% CI=1.01-1.07, P=0.02). We did not observe a difference in these relationships when stratified by breast tumor estrogen receptor status. We repeated this analysis using genetic variants independent of the leading association at core HDL pathway genes and found that these variants were also associated with risk for breast cancers (OR=1.11, 95% CI=1.06-1.16, P=1.5x10-6), including gene-specific associations at ABCA1, APOE-APOC1-APOC4-APOC2 and CETP. In addition, we find evidence that genetic variation at the ABO locus affects both lipid levels and breast cancer.\n\nConclusionsGenetically elevated plasma HDL levels appear to increase breast cancer risk. Future studies are required to understand the mechanism underlying this putative causal relationship, with the goal to develop potential therapeutic strategies aimed at altering the HDL-mediated effect on breast cancer risk.

genomics