bioRxiv ScienceSearch

Biology subjects

Visscher, P. M.

Publications and source records attributed to Visscher, P. M..

At least 19 recordsLinked to original sources

Bayesian reassessment of the epigenetic architecture of complex traits

1Epigenetic DNA modification is partly under genetic control, and occurs in response to a wide range of environmental exposures. Linking epigenetic marks to clinical outcomes may provide greater insight into underlying molecular processes of disease, assist in the identification of therapeutic targets, and improve risk prediction. Here, we present a statistical approach, based on Bayesian inference, that estimates associations between disease risk and all measured epigenetic probes jointly, automatically controlling for both data structure (including cell-count effects, relatedness, and experimental batch effects) and correlations among probes. We benchmark our approach in simulation study, finding improved estimation of probe associations across a wide range of scenarios over existing approaches. Our method estimates the total proportion of disease risk captured by epigenetic probe variation, and when we applied it to measures of body mass index (BMI) and cigarette consumption behaviour in 5,101 individuals, we find that 66.7% (95% CI 60.0-72.8) of the variation in BMI and 67.7% (95% CI 58.4-76.9) of the variation in cigarette consumption can be captured by methylation array data from whole blood, independent of the variation explained by single nucleotide polymorphism markers. We find novel associations, with smoking behaviour associated with a methylation probe at the MNDA gene with >95% posterior inclusion probability, which is a myeloid cell nuclear differentiation antigen gene previously implicated as a biomarker for inflammation and non-Hodgkin lymphoma risk. We conduct unique genome-wide enrichment analyses, identifying blood cholesterol, lipid transport and sterol metabolism pathways for BMI, and response to xenobiotic stimulus and negative regulation of RNA polymerase II promoter transcription for smoking, all with >95% posterior inclusion probability of having methylation probes with associations >1.5 times larger than the average. Finally, we improve phenotypic prediction in two independent cohorts by 28.7% and 10.2% for BMI and smoking respectively over a LASSO model. These results imply that probe measures may capture large amounts of variance because they are likely a consequence of the phenotype rather than a cause. As a result, 'omics' data may enable accurate characterization of disease progression and identification of individuals who are on a path to disease. Our approach facilitates better understanding of the underlying epigenetic architecture of complex common disease and is applicable to any kind of genomics data.

genomics

OSCA: a tool for omic-data-based complex trait analysis

The rapid increase of omic data in the past decades has greatly facilitated the investigation of associations between omic profiles such as DNA methylation (DNAm) and complex traits in large cohorts. Here, we proposed a mixed-linear-model-based method (called MOMENT) that tests for association between a DNAm probe and trait with all other distal probes fitted in multiple random-effect components to account for the effects of unobserved confounders as well as the correlations between distal probes induced by the confounders. We demonstrated by simulations that MOMENT showed a lower false positive rate and more robustness than existing methods. MOMENT has been implemented in a versatile software package (called OSCA) together with a number of other implementations for omic-data-based analysis including the estimation of variance in a trait captured by all measures of multiple omic profiles, omic-data-based quantitative trait locus (xQTL) analysis, and meta-analysis of xQTL data.

bioinformatics

Parkinson disease age of onset GWAS: defining heritability, genetic loci and a-synuclein mechanisms

Increasing evidence supports an extensive and complex genetic contribution to Parkinsons disease (PD). Previous genome-wide association studies (GWAS) have shed light on the genetic basis of risk for this disease. However, the genetic determinants of PD age of onset are largely unknown. Here we performed an age of onset GWAS based on 28,568 PD cases. We estimated that the heritability of PD age of onset due to common genetic variation was ~0.11, lower than the overall heritability of risk for PD (~0.27) likely in part because of the subjective nature of this measure. We found two genome-wide significant association signals, one at SNCA and the other a protein-coding variant in TMEM175, both of which are known PD risk loci and a Bonferroni corrected significant effect at other known PD risk loci, INPP5F/BAG3, FAM47E/SCARB2, and MCCC1. In addition, we identified that GBA coding variant carriers had an earlier age of onset compared to non-carriers. Notably, SNCA, TMEM175, SCARB2, BAG3 and GBA have all been shown to either directly influence alpha-synuclein aggregation or are implicated in alpha-synuclein aggregation pathways. Remarkably, other well-established PD risk loci such as GCH1, MAPT and RAB7L1/NUCKS1 (PARK16) did not show a significant effect on age of onset of PD. While for some loci, this may be a measure of power, this is clearly not the case for the MAPT locus; thus genetic variability at this locus influences whether but not when an individual develops disease. We believe this is an important mechanistic and therapeutic distinction. Furthermore, these data support a model in which alpha-synuclein and lysosomal mechanisms impact not only PD risk but also age of disease onset and highlights that therapies that target alpha-synuclein aggregation are more likely to be disease-modifying than therapies targeting other pathways.

genetics

The effect of X-linked dosage compensation on complex trait variation

Quantitative genetics theory predicts that X-chromosome dosage compensation between sexes will have a detectable effect on the amount of genetic and therefore phenotypic trait variances at associated loci in males and females. Here, we systematically examine the role of dosage compensation in complex trait variation in humans in 20 complex traits in a sample of more than 450,000 individuals from the UK Biobank and in 1,600 gene expression traits from a sample of 2,000 individuals as well as across-tissue gene expression from the GTEx resource. We find, on average, twice as much genetic variation for complex traits due to X-linked loci in males compared to females, consistent with a negligible effect of predicted escape from X-inactivation on complex trait variation across traits and also detect biologically relevant X-linked heterogeneity between the sexes for a number of complex traits.

genetics

Epigenetic signatures of starting and stopping smoking

BackgroundMultiple studies have made robust associations between differential DNA methylation and exposure to cigarette smoke. But whether a DNA methylation phenotype is established immediately upon exposure, or only after prolonged exposure is less well-established. Here, we assess DNA methylation patterns in current smokers in response to dose and duration of exposure, along with the effects of smoking cessation on DNA methylation in former smokers.\n\nMethodsDimensionality reduction was applied to DNA methylation data at 90 previously identified smoking-associated CpG sites for over 4,900 individuals in the Generation Scotland cohort. K-means clustering was performed to identify clusters associated with current and never smoker status based on these methylation patterns. Cluster assignments were assessed with respect to duration of exposure in current smokers (years as a smoker), time since smoking cessation in former smokers (years), and dose (cigarettes per day).\n\nResultsTwo clusters were specified, corresponding to never smokers (97.5% of whom were assigned to Cluster 1) and current smokers (81.1% of whom were assigned to Cluster 2). The exposure time point from which >50% of current smokers were assigned to the smoker-enriched cluster varied between 5-9 years in heavier smokers and between 15-19 years in lighter smokers. Low-dose former smokers were more likely to be assigned to the never smoker-enriched cluster from the first year following cessation. In contrast, a period of at least two years was required before the majority of former high-dose smokers were assigned to the never smoker-enriched cluster.\n\nConclusionsOur findings suggest that smoking-associated DNA methylation changes are a result of prolonged exposure to cigarette smoke, and can be reversed following cessation. The length of time in which these signatures are established and recovered is dose dependent. Should DNA methylation-based signatures of smoking status be predictive of smoking-related health outcomes, our findings may provide an additional criterion on which to stratify risk.

genomics

Expectation of the intercept from bivariate LD score regression in the presence of population stratification

Linkage disequilibrium (LD) score regression is an increasingly popular method used to quantify the level of confounding in genome-wide association studies (GWAS) or to estimate heritability and genetic correlation between traits. When applied to a pair of GWAS, the LD score regression (LDSC) methodology produces a statistic, referred to as the bivariate LDSC intercept, which deviation from 0 is classically interpreted as an indication of sample overlap between the two GWAS. Here we propose an extension of the theory underlying the bivariate LDSC methodology, which accounts for population stratification within and between GWAS. Our extended theory predicts an inflation of the bivariate LDSC intercept when sample sizes and heritability are large, even in the absence of sample overlap. We illustrate our theoretical results with simulations based on actual SNP genotypes and we propose a re-interpretation of previously published results in the light of our extended theory.

genetics

Meta-analysis of genome-wide association studies for body fat distribution in 694,649 individuals of European ancestry

One in four adults worldwide are either overweight or obese. Epidemiological studies indicate that the location and distribution of excess fat, rather than general adiposity, is most informative for predicting risk of obesity sequellae, including cardiometabolic disease and cancer. We performed a genome-wide association study meta-analysis of body fat distribution, measured by waist-to-hip ratio adjusted for BMI (WHRadjBMI), and identified 463 signals in 346 loci. Heritability and variant effects were generally stronger in women than men, and we found approximately one-third of all signals to be sexually dimorphic. The 5% of individuals carrying the most WHRadjBMI-increasing alleles were 1.62 times more likely than the bottom 5% to have a WHR above the thresholds used for metabolic syndrome. These data, made publicly available, will inform the biology of body fat distribution and its relationship with disease.

genetics

Imprint of Assortative Mating on the Human Genome

Non-random mate-choice with respect to complex traits is widely observed in humans, but whether this reflects true phenotypic assortment, environment (social homogamy) or convergence after choosing a partner is not known. Understanding the causes of mate choice is important, because assortative mating (AM) if based upon heritable traits, has genetic and evolutionary consequences. AM is predicted under Fishers classical theory1 to induce a signature in the genome at trait-associated loci that can be detected and quantified. Here, we develop and apply a method to quantify AM on a specific trait by estimating the correlation ({theta}) between genetic predictors of the trait from SNPs on odd versus even chromosomes. We show by theory and simulation that the effect of AM can be distinguished from population stratification. We applied this approach to 32 complex traits and diseases using SNP data from [~]400,000 unrelated individuals of European ancestry. We found significant evidence of AM for height ({theta}=3.2%) and educational attainment ({theta}=2.7%), both consistent with theoretical predictions. Overall, our results imply that AM involves multiple traits, affects the genomic architecture of loci that are associated with these traits and that the consequence of mate choice can be detected from a random sample of genomes.

genetics

Epigenetic prediction of complex traits and death

BackgroundGenome-wide DNA methylation (DNAm) profiling has allowed for the development of molecular predictors for a multitude of traits and diseases. Such predictors may be more accurate than the self-reported phenotypes, and could have clinical applications. Here, penalised regression models were used to develop DNAm predictors for body mass index (BMI), smoking status, alcohol consumption, and educational attainment in a cohort of 5,100 individuals. Using an independent test cohort comprising 906 individuals, the proportion of phenotypic variance explained in each trait was examined for DNAm-based and genetic predictors. Receiver operator characteristic curves were generated to investigate the predictive performance of DNAm-based predictors, using dichotomised phenotypes. The relationship between DNAm scores and all-cause mortality (n = 214 events) was assessed via Cox proportional-hazards models.\n\nResultsThe DNAm-based predictors explained different proportions of the phenotypic variance for BMI (12%), smoking (60%), alcohol consumption (12%) and education (3%). The combined genetic and DNAm predictors explained 20% of the variance in BMI, 61% in smoking, 13% in alcohol consumption, and 6% in education. DNAm predictors for smoking, alcohol, and education but not BMI predicted mortality in univariate models. The predictors showed moderate discrimination of obesity (AUC=0.67) and alcohol consumption (AUC=0.75), and excellent discrimination of current smoking status (AUC=0.98). There was poorer discrimination of college-educated individuals (AUC=0.59).\n\nConclusionsDNAm predictors correlate with lifestyle factors that are associated with health and mortality. They may supplement DNAm-based predictors of age to identify the lifestyle profiles of individuals and predict disease risk.\n\nList of abbreviations

genomics

Novel susceptibility loci and genetic regulation mechanisms for type 2 diabetes

We conducted a meta-analysis of genome-wide association studies (GWAS) with [~]16 million genotyped/imputed genetic variants in 62,892 type 2 diabetes (T2D) cases and 596,424 controls of European ancestry. We identified 139 common and 4 rare (minor allele frequency < 0.01) variants associated with T2D, 42 of which (39 common and 3 rare variants) were independent of the known variants. Integration of the gene expression data from blood (n = 14,115 and 2,765) and other T2D-relevant tissues (n = up to 385) with the GWAS results identified 33 putative functional genes for T2D, three of which were targeted by approved drugs. A further integration of DNA methylation (n = 1,980) and epigenomic annotations data highlighted three putative T2D genes (CAMK1D, TP53INP1 and ATP5G1) with plausible regulatory mechanisms whereby a genetic variant exerts an effect on T2D through epigenetic regulation of gene expression. We further found evidence that the T2D-associated loci have been under purifying selection.

genetics

Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood

Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes associated with brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top associated cis-expression (cis-eQTLs or cis-mQTLs) between brain and blood for genes expressed (or CpG sites methylated) in both tissues, while accounting for errors in their estimated effects (rb). Using publicly available data (n = 72 to l,366), we find that the genetic effects of cis-eQTLs (PeQTL < 5x10-8) or mQTLs (PmQTL < 1x10-10) are highly correlated between independent brain and blood samples ([Formula] with SE = 0.015 for cis-eQTL and [Formula] with SE = 0.006 for cis-mQTLs). Using meta-analyzed brain eQTL/mQTL data (n = 526 to 1,194), we identify 61 genes and 167 DNA methylation (DNAm) sites associated with 4 brain-related traits and disorders. Most of these associations are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = l,980 to 14,115). We further find that cis-eQTLs with tissue-specific effects are approximately uniformly distributed across all the functional annotation categories, and that mean difference in gene expression level between brain and blood is almost independent of the difference in the corresponding cis-eQTL effect. Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-eQTL or cis-mQTL data with large sample sizes.

genetics

Meta-analysis of genome-wide association studies for height and body mass index in ~700,000 individuals of European ancestry

Genome-wide association studies (GWAS) stand as powerful experimental designs for identifying DNA variants associated with complex traits and diseases. In the past decade, both the number of such studies and their sample sizes have increased dramatically. Recent GWAS of height and body mass index (BMI) in [~]250,000 European participants have led to the discovery of [~]700 and [~]100 nearly independent SNPs associated with these traits, respectively. Here we combine summary statistics from those two studies with GWAS of height and BMI performed in [~]450,000 UK Biobank participants of European ancestry. Overall, our combined GWAS meta-analysis reaches N[~]700,000 individuals and substantially increases the number of GWAS signals associated with these traits. We identified 3,290 and 716 near-independent SNPs associated with height and BMI, respectively (at a revised genome-wide significance threshold of p<1 x 10-8), including 1,185 height-associated SNPs and 554 BMI-associated SNPs located within loci not previously identified by these two GWAS. The genome-wide significant SNPs explain [~]24.6% of the variance of height and [~]5% of the variance of BMI in an independent sample from the Health and Retirement Study (HRS). Correlations between polygenic scores based upon these SNPs with actual height and BMI in HRS participants were 0.44 and 0.20, respectively. From analyses of integrating GWAS and eQTL data by Summary-data based Mendelian Randomization (SMR), we identified an enrichment of eQTLs amongst lead height and BMI signals, prioritisting 684 and 134 genes, respectively. Our study demonstrates that, as previously predicted, increasing GWAS sample sizes continues to deliver, by discovery of new loci, increasing prediction accuracy and providing additional data to achieve deeper insight into complex trait biology. All summary statistics are made available for follow up studies.

genetics

GWAS on family history of Alzheimer’s disease

Alzheimers disease (AD) is a public health priority for the 21st century. Risk reduction currently revolves around lifestyle changes with much research trying to elucidate the biological underpinnings. Using self-report of parental history of Alzheimers dementia for case ascertainment in a genome-wide association study of over 300,000 participants from UK Biobank (32,222 maternal cases, 16,613 paternal cases) and meta-analysing with published consortium data (n=74,046 with 25,580 cases across the discovery and replication analyses), six new AD-associated loci (P<5x10-8) are identified. Three contain genes relevant for AD and neurodegeneration: ADAM10, ADAMTS4, and ACE. Suggestive loci include drug targets such as VKORC1 (warfarin dose) and BZRAP1 (benzodiazepine receptor). We report evidence that association of SNPs and AD at the PVR gene is potentially mediated by both gene expression and DNA methylation in the prefrontal cortex. Our discovered loci may help to elucidate the biological mechanisms underlying AD and, given that many are existing drug targets for other diseases and disorders, warrant further exploration for potential precision medicine applications.

genetics

Epigenetic influences on aging: a longitudinal genome-wide methylation study in old Swedish twins

Age-related changes in DNA methylation have been observed in many cross-sectional studies, but longitudinal evidence is still very limited. Here, we aimed to characterize longitudinal age-related methylation patterns (Illumina HumanMethylation450 array) using 1011 blood samples collected from 385 old Swedish twins (mean age of 69 at baseline) up to five times over 20 years. We identified 1316 age-associated methylation sites (p<1.3x10-7) using a longitudinal epigenome-wide association study design. We measured how estimated cellular compositions changed with age and how much they confounded the age effect. We validated the results in two independent longitudinal cohorts, where 118 CpGs were replicated in PIVUS (p<3.9x10-5) and 594 were replicated in LBC (p<5.1x10-5). Functional annotation of age-associated CpGs showed enrichment in CCCTC-binding factor (CTCF) and other unannotated transcription factor binding sites. We further investigated genetic influences on methylation (methylation quantitative trait loci) and found no interaction between age and genetic effects in the 1316 age-associated CpGs. Moreover, in the same CpGs, methylation differences within twin pairs increased over time, where monozygotic twins had smaller intra-pair differences than dizygotic twins. We show that age-related methylation changes persist in a longitudinal perspective, and are fairly stable across cohorts. Moreover, the changes are under genetic influence, although this effect is independent of age. In addition, inter-individual methylation variations increase over time, especially in age-associated CpGs, indicating the increase of environmental contributions on DNA methylation with age.

genetics

Equivalence of LD-Score Regression and Individual-Level-Data Methods

LD-score (LDSC) regression disentangles the contribution of polygenic signal, in terms of SNP-based heritability, and population stratification, in terms of a so-called intercept, to GWAS test statistics. Whereas LDSC regression uses summary statistics, methods like Haseman-Elston (HE) regression and genomic-relatedness-matrix (GRM) restricted maximum likelihood infer parameters such as SNP-based heritability from individual-level data directly. Therefore, these two types of methods are typically considered to be profoundly different. Nevertheless, recent work has revealed that LDSC and HE regression yield near-identical SNP-based heritability estimates when confounding stratification is absent. We now extend the equivalence; under the stratification assumed by LDSC regression, we show that the intercept can be estimated from individual-level data by transforming the coefficients of a regression of the phenotype on the leading principal components from the GRM. Using simulations, considering various degrees and forms of population stratification, we find that intercept estimates obtained from individual-level data are nearly equivalent to estimates from LDSC regression (R2 > 99%). An empirical application corroborates these findings. Hence, LDSC regression is not profoundly different from methods using individual-level data; parameters that are identified by LDSC regression are also identified by methods using individual-level data. In addition, our results indicate that, under strong stratification, there is misattribution of stratification to the slope of LDSC regression, inflating estimates of SNP-based heritability from LDSC regression ceteris paribus. Hence, the intercept is not a panacea for population stratification. Consequently, LDSC-regression estimates should be interpreted with caution, especially when the intercept estimate is significantly greater than one.

genetics

Identification of 55,000 Replicated DNA Methylation QTL

DNA methylation plays an important role in the regulation of transcription. Genetic control of DNA methylation is a potential candidate for explaining the many identified SNP associations with disease that are not found in coding regions. We replicated 52,916 cis and 2,025 trans DNA methylation quantitative trait loci (mQTL) using methylation measured on Illumina HumanMethylation450 arrays in the Brisbane Systems Genetics Study (n=614 from 177 families) and the Lothian Birth Cohorts of 1921 and 1936 (combined n = 1366). The trans mQTL SNPs were found to be over-represented in 1Mbp subtelomeric regions, and on chromosomes 16 and 19. There was a significant increase in trans mQTL DNA methylation sites in upstream and 5 UTR regions. No association was observed between either the SNPs or DNA methylation sites of trans mQTL and telomere length. The genetic heritability of a number of complex traits and diseases was partitioned into components due to mQTL and the remainder of the genome. Significant enrichment was observed for height (p = 2.1x10-10), ulcerative colitis (p = 2x10-5), Crohns disease (p = 6x10-8) and coronary artery disease (p = 5.5x10-6) when compared to a random sample of SNPs with matched minor allele frequency, although this enrichment is explained by the genomic location of the mQTL SNPs.

genomics

GWAS of epigenetic ageing rates in blood reveals a critical role for TERT

DNA methylation age is an accurate biomarker of chronological age and predicts lifespan, but its underlying molecular mechanisms are unknown. In this genome-wide association study of 9,907 individuals, we found gene variants mapping to five loci associated with intrinsic epigenetic age acceleration (IEAA) and gene variants in 3 loci associated extrinsic epigenetic age acceleration (EEAA). Mendelian randomization analysis suggested causal influences of menarche and menopause on IEAA and lipid levels on IEAA and EEAA. Variants associated with longer leukocyte telomere length (LTL) in the telomerase reverse transcriptase gene (TERT) locus at 5p15.33 confer higher IEAA (P<2.7x10-11). Causal modelling indicates TERT-specific and independent effects on LTL and IEAA. Experimental hTERT expression in primary human fibroblasts engenders a linear increase in DNA methylation age with cell population doubling number. Together, these findings indicate a critical role for hTERT in regulating the DNA methylation clock, in addition to its established role of compensating for cell replication-dependent telomere shortening.

genetics

MTAG: Multi-Trait Analysis of GWAS

We introduce Multi-Trait Analysis of GWAS (MTAG), a method for joint analysis of summary statistics from GWASs of different traits, possibly from overlapping samples. We apply MTAG to summary statistics for depressive symptoms (Neff = 354,862), neuroticism (N = 168,105), and subjective well-being (N = 388,538). Compared to 32, 9, and 13 genome-wide significant loci in the single-trait GWASs (most of which are themselves novel), MTAG increases the number of loci to 64, 37, and 49, respectively. Moreover, association statistics from MTAG yield more informative bioinformatics analyses and increase variance explained by polygenic scores by approximately 25%, matching theoretical expectations.

genomics