bioRxiv ScienceSearch

Biology subjects

Davey Smith, G.

Publications and source records attributed to Davey Smith, G..

At least 55 records · Page 3Linked to original sources

Cigarette smoking and personality: Investigating causality using Mendelian randomization

BackgroundDespite the well-documented association between smoking and personality traits such as neuroticism and extraversion, little is known about the potential causal nature of these findings. If it were possible to unpick the association between personality and smoking, it may be possible to develop more targeted smoking cessation programmes that could lead to both improved uptake and efficacy.\n\nMethodsRecent genome-wide association studies (GWAS) have identified variants robustly associated with both smoking phenotypes and personality traits. Here we use publicly available GWAS summary statistics in addition to data from UK Biobank to investigate the link between smoking and personality. We first estimated genetic overlap between traits using LD score regression and then applied both one- and two-sample Mendelian randomization methods to unpick the nature of this relationship.\n\nResultsWe found clear evidence of a modest genetic correlation between smoking behaviours and both neuroticism and extraversion, suggesting shared genetic aetiology. We found some evidence to suggest an association between neuroticism and increased smoking initiation. We also found some evidence that personality traits appear to be causally linked to certain smoking phenotypes: higher neuroticism and heavier cigarette consumption, and higher extraversion and increased odds of smoking initiation. The latter finding could lead to more targeted smoking prevention programmes.\n\nConclusionThe association between neuroticism and cigarette consumption lends support to the self-medication hypothesis, while the association between extraversion and smoking initiation could lead to more targeted smoking prevention programmes.

epidemiology

Searching for the causal effects of BMI in over 300 000 individuals, using Mendelian randomization

Mendelian randomization (MR) has been used to estimate the causal effect of body mass index (BMI) on particular traits thought to be affected by BMI. However, BMI may also be a modifiable, causal risk factor for outcomes where there is no prior reason to suggest that a causal effect exists. We perform a MR phenome-wide association study (MR-pheWAS) to search for the causal effects of BMI in UK Biobank (n=334 968), using the PHESANT open-source phenome scan tool. Of the 20 461 tests performed, our MR-pheWAS identified 519 associations below a stringent P value threshold corresponding to a 5% estimated false discovery rate, including many previously identified causal effects. We also identified several novel effects, including protective effects of higher BMI on a set of psychosocial traits, identified initially in our preliminary MR-pheWAS and replicated in an independent subset of UK Biobank. Such associations need replicating in an independent sample.

epidemiology

Examining the genetic influences of educational attainment and the validity of value-added measures of progress

In this study, we estimate (i) the SNP heritability of educational attainment at three time points throughout the compulsory educational lifecourse; (ii) the SNP heritability of value-added measures of educational progress built from test data; and (iii) the extent to which value-added measures built from teacher rated ability may be biased due to measurement error. We utilise a genome wide approach using generalized restricted maximum likelihood (GCTA-GREML) to determine the total phenotypic variance in educational attainment and value-added measures that is attributable to common genetic variation across the genome within a sample of unrelated individuals from a UK birth cohort, the Avon Longitudinal Study of Parents and Children. Our findings suggest that the heritability of educational attainment measured using point score test data increases with age from 47% at age 11 to 61% at age 16. We also find that genetic variation does not contribute towards value-added measures created only from educational attainment point score data, but it does contribute a small amount to measures that additionally control for background characteristics (up to 20.09% [95%CI: 6.06 to 35.71] from age 11 to 14). Finally, our results show that value-added measures built from teacher rated ability have higher heritability than those built from exam scores. Our findings suggest that the heritability of educational attainment increases through childhood and adolescence. Value-added measures based upon fine grain point scores may be less prone to between-individual genomic differences than measures that control for students backgrounds, or those built from more subjective measures such as teacher rated ability.

genetics

Conditioning on a collider may induce spurious associations: Do the results of Gale et al. (2017) support a protective effect of neuroticism in population sub-groups?

Introduction Introduction Methods Results Discussion References Gale and colleagues (Gale et al., 2017) examined the association between neuroticism and mortality in a large sample (N > 300,000) drawn from the UK Biobank study (Sudlow et al., 2015). They observed that neuroticism was associated with an increase in all-cause mortality, but that following adjustment for self-rated health neuroticism was associated with a reduction in all-cause mortality. Further analyses stratified on self-rated health suggested that higher neuroticism was associated with reduced mortality only among those with fair or poor self-rated health. The authors conclude that neuroticism may have protective effects among certain sub-groups, and finding that generated substantial interest (TIME, 2017).\n\nThe availability ...

epidemiology

The molecular genetics of participation in the Avon Longitudinal Study of Parents and Children

BackgroundIt is often assumed that selection (including participation and dropout) does not represent an important source of bias in genetic studies. However, there is little evidence to date on the effect of genetic factors on participation.\n\nMethodsUsing data on mothers (N=7,486) and children (N=7,508) from the Avon Longitudinal Study of Parents and Children, we 1) examined the association of polygenic risk scores for a range of socio-demographic, lifestyle characteristics and health conditions related to continued participation, 2) investigated whether associations of polygenic scores with body mass index (BMI; derived from self-reported weight and height) and self-reported smoking differed in the largest sample with genetic data and a sub-sample who participated in a recent follow-up and 3) determined the proportion of variation in participation explained by common genetic variants using genome-wide data.\n\nResultsWe found evidence that polygenic scores for higher education, agreeableness and openness were associated with higher participation and polygenic scores for smoking initiation, higher BMI, neuroticism, schizophrenia, ADHD and depression were associated with lower participation. Associations between the polygenic score for education and self-reported smoking differed between the largest sample with genetic data (OR for ever smoking per SD increase in polygenic score:0.85, 95% CI:0.81,0.89) and sub-sample (OR:0.95, 95% CI:0.88,1.02). In genome-wide analysis, single nucleotide polymorphism based heritability explained 17-31% of variability in participation.\n\nConclusionsGenetic association studies, including Mendelian randomization, can be biased by selection, including loss to follow-up. Genetic risk for dropout should be considered in all analyses of studies with selective participation.

genetics

Improving the visualisation, interpretation and analysis of two-sample summary data Mendelian randomization via the radial plot and radial regression

BackgroundSummary data furnishing a two-sample Mendelian randomization study are often visualized with the aid of a scatter plot, in which single nucleotide polymorphism (SNP)-outcome associations are plotted against the SNP-exposure associations to provide an immediate picture of the causal effect estimate for each individual variant. It is also convenient to overlay the standard inverse variance weighted (IVW) estimate of causal effect as a fitted slope, to see whether an individual SNP provides evidence that supports, or conflicts with, the overall consensus. Unfortunately, the traditional scatter plot is not the most appropriate means to achieve this aim whenever SNP-outcome associations are estimated with varying degrees of precision and this is reflected in the analysis.\n\nMethodsWe propose instead to use a small modification of the scatter plot - the Galbraith radial plot - for the presentation of data and results from an MR study, which enjoys many advantages over the original method. On a practical level it removes the need to recode the genetic data and enables a more straightforward detection of outliers and influential data points. Its use extends beyond the purely aesthetic, however, to suggest a more general modelling framework to operate within when conducting an MR study, including a new form of MR-Egger regression.\n\nResultsWe illustrate the methods using data from a two-sample Mendelian randomization study to probe the causal effect of systolic blood pressure on coronary heart disease risk, allowing for the possible effects of pleiotropy. The radial plot is shown to aid the detection of a single outlying variant which is responsible for large differences between IVW and MR-Egger regression estimates. Several additional plots are also proposed for informative data visualisation.\n\nConclusionThe radial plot should be considered in place of the scatter plot for visualising, analysing and interpreting data from a two-sample summary data MR study. Software is provided to help facilitate its use.

epidemiology

Selection bias in instrumental variable analyses

Participants in epidemiological and genetic studies are rarely true random samples of the populations they are intended to represent, and both known and unknown factors can influence participation in a study (known as selection into a study). The circumstances in which selection causes bias in an instrumental variable (IV) analysis are not widely understood by practitioners of IV analyses. We use directed acyclic graphs (DAGs) to depict assumptions about the selection mechanism (factors affecting selection) and show how DAGs can be used to determine when a two-stage least squares (2SLS) IV analysis is biased by different selection mechanisms. Via simulations, we show that selection can result in a biased IV estimate with substantial confidence interval undercoverage, and the level of bias can differ between instrument strengths, a linear and nonlinear exposure-instrument association, and a causal and noncausal exposure effect. We present an application from the UK Biobank study, which is known to be a selected sample of the general population. Of interest was the causal effect of education on the decision to smoke. The 2SLS exposure estimates were very different between the IV analysis ignoring selection and the IV analysis which adjusted for selection (e.g., 1.8 [95% confidence interval -1.5, 5.0] and -4.5 [-6.6, -2.4], respectively). We conclude that selection bias can have a major effect on an IV analysis and that statistical methods for estimating causal effects using data from nonrandom samples are needed.

epidemiology

Detecting and correcting for bias in Mendelian randomization analyses using gene-by-environment interactions

BackgroundMendelian randomization has developed into an established method for strengthening causal inference and estimating causal effects, largely due to the proliferation of genome-wide association studies. However, genetic instruments remain controversial as pleiotropic effects can introduce bias into causal estimates. Recent work has highlighted the potential of gene-environment interactions in detecting and correcting for pleiotropic bias in Mendelian randomization analyses.\n\nMethodsWe introduce MR using Gene-by-Environment interactions (MRGxE) as a framework capable of identifying and correcting for pleiotropic bias, drawing upon developments in econometrics and epidemiology. If an instrument-covariate interaction induces variation in the association between a genetic instrument and exposure, it is possible to identify and correct for pleiotropic effects. The interpretation of MRGxE is similar to conventional summary Mendelian randomization approaches, with a particular advantage of MRGxE being the ability to assess the validity of an individual instrument.\n\nResultsWe investigate the effect of BMI upon systolic blood pressure (SBP) using data from the UK Biobank and the GIANT consortium using a single instrument (a weighted allelic score). We find MRGxE produces findings in agreement with MR Egger regression in a two-sample summary MR setting, however, association estimates obtained across all methods differ considerably when excluding related participants or individuals of non-European ancestry. This could be a consequence of selection bias, though there is also potential for introducing bias by using a mixed ancestry population. Further, we assess the performance of MRGxE with respect to identifying and correcting for horizontal pleiotropy in a simulation setting, highlighting the utility of the approach even when the MRGxE assumptions are violated.\n\nConclusionsBy utilising instrument-covariate interactions within a linear regression framework, it is possible to identify and correct for pleiotropic bias, provided the average magnitude of pleiotropy is constant across interaction covariate subgroups.\n\nO_TEXTBOXKey MessagesO_LIInstrument-covariate interactions can be used to identify pleiotropic bias in Mendelian randomization analyses, provided they induce sufficient variation in the association between the genetic instrument and exposure.\nC_LIO_LIBy regressing the gene-outcome association upon the gene-exposure association across interaction covariate subgroups, it is possible to obtain an estimate of the average pleiotropic effect and a causal effect estimate.\nC_LIO_LIThe interpretation of MRGxE is analogous to that of MR-Egger regression.\nC_LIO_LIThe approach serves as a valuable test for directional pleiotropy and can be used to inform instrument selection.\nC_LI\n\nC_TEXTBOX

epidemiology

Systematic Mendelian randomization framework elucidates hundreds of genetic loci which may influence disease through changes in DNA methylation levels

We have undertaken an extensive Mendelian randomization (MR) study using methylation quantitative trait loci (mQTL) as genetic instruments to assess the potential causal relationship between genetic variation, DNA methylation and 139 complex traits. Using two-sample MR, we observed 1,191 effects across 62 traits where genetic variants were associated with both proximal DNA methylation (i.e. cis-mQTL) and complex trait variation (P<1.39x10-08). Joint likelihood mapping provided evidence that the causal mQTL for 364 of these effects across 58 traits was also likely the causal variant for trait variation. These effects showed a high rate of replication in the UK Biobank dataset for 14 selected traits, as 121 of the attempted 129 effects replicated. Integrating expression quantitative trait loci (eQTL) data suggested that genetic variants responsible for 319 of the 364 mQTL effects also influence gene expression, which indicates a coordinated system of effects that are consistent with causality. CpG sites were enriched for histone mark peaks in tissue types relevant to their associated trait and implicated genes were enriched across relevant biological pathways. Though we are unable to distinguish mediation from horizontal pleiotropy in these analyses, our findings should prove valuable in identifying candidate loci for further evaluation and help develop mechanistic insight into the aetiology of complex disease.

genetics

Investigating causality in associations between education and smoking: A two-sample Mendelian randomization study

BackgroundLower educational attainment is associated with increased rates of smoking, but ascertaining causality is challenging. We used two-sample Mendelian randomization (MR) analyses of summary statistics to examine whether educational attainment is causally related to smoking.\n\nMethods and FindingsWe used summary statistics from genome-wide association studies of educational attainment and a range of smoking phenotypes (smoking initiation, cigarettes per day, cotinine levels and smoking cessation). Various complementary MR techniques (inverse-variance weighted regression, MR Egger, weighted-median regression) were used to test the robustness of our results. We found broadly consistent evidence across these techniques that higher educational attainment leads to reduced likelihood of smoking initiation, reduced heaviness of smoking among smokers (as measured via self-report and cotinine levels), and greater likelihood of smoking cessation among smokers.\n\nConclusionsOur findings indicate a causal association between low educational attainment and increased risk of smoking, and may explain the observational associations between educational attainment and adverse health outcomes such as risk of coronary heart disease.

epidemiology

Effect modification of FADS2 polymorphisms on the association between breastfeeding and intelligence: results from a collaborative meta-analysis

BackgroundAccumulating evidence suggests that breastfeeding benefits the childrens intelligence. Long-chain polyunsaturated fatty acids (LC-PUFAs) present in breast milk may explain part of this association. Under a nutritional adequacy hypothesis, an interaction between breastfeeding and genetic variants associated with endogenous LC-PUFAs synthesis might be expected. However, the literature on this topic is controversial.\n\nMethods and FindingsWe investigated this GenexEnvironment interaction in a de novo meta-analysis involving >12,000 individuals in the primary analysis, and >45,000 individuals in a secondary analysis using relaxed inclusion criteria. Our primary analysis used ever breastfeeding, FADS2 polymorphisms rs174575 and rs1535 coded assuming a recessive effect of the G allele, and intelligence quotient (IQ) in Z scores. Using random effects meta-analysis, ever breastfeeding was associated with 0.17 (95% CI: 0.03; 0.32) higher Z scores in IQ, or about 2.1 points. There was no strong evidence of interaction, with pooled covariate-adjusted interaction coefficients (i.e., difference between genetic groups of the difference in IQZ scores comparing ever with never breastfed individuals) of 0.12 (95% CI: -0.19; 0.43) and 0.06 (95% CI: -0.16; 0.27) for the rs174575 and rs1535 variants, respectively. Secondary analyses corroborated these results. In studies with >5.85 and <5.85 months of breastfeeding duration, pooled estimates for the rs174575 variant were 0.50 (95% CI: -0.06; 1.06) and 0.14 (95% CI: -0.10; 0.38), respectively, and 0.27 (95% CI: -0.28; 0.82) and -0.01 (95% CI: -0.19; 0.16) for the rs1535 variant. However, between-group comparisons were underpowered.\n\nConclusionsOur findings do not support an interaction between ever breastfeeding and FADS2 polymorphisms. However, our subgroup analysis raises the possibility that breastfeeding supplies LC-PUFAs requirements for cognitive development (if such threshold exists) if it lasts for some (currently unknown) time. Future studies in large individual-level datasets would allow properly powered subgroup analyses and would improve our understanding on the role of breastfeeding duration in the breastfeedingxFADS2 interaction.

epidemiology

Developmental changes within the genetic architecture of social communication behaviour: A multivariate study of genetic variance in unrelated individuals

BackgroundRecent analyses of trait-disorder overlap suggest that psychiatric dimensions may relate to distinct sets of genes that exert their maximum influence during different periods of development. This includes analyses of social-communciation difficulties that share, depending on their developmental stage, stronger genetic links with either Autism Spectrum Disorder or schizophrenia. Here we developed a multivariate analysis framework in unrelated individuals to model directly the developmental profile of genetic influences contributing to complex traits, such as social-communication difficulties, during a [~]10-year period spanning childhood and adolescence.\n\nMethodsLongitudinally assessed quantitative social-communication problems (N[&le;] 5,551) were studied in participants from a UK birth cohort (ALSPAC, 8 to 17 years). Using standardised measures, genetic architectures were investigated with novel multivariate genetic-relationship-matrix structural equation models (GSEM) incorporating whole-genome genotyping information. Analogous to twin research, GSEM included Cholesky decomposition, common pathway and independent pathway models.\n\nResultsA 2-factor Cholesky decomposition model described the data best. One genetic factor was common to SCDC measures across development, the other accounted for independent variation at 11 years and later, consistent with distinct developmental profiles in trait-disorder overlap. Importantly, genetic factors operating at 8 years explained only [~]50% of the genetic variation at 17 years.\n\nConclusionUsing latent factor models, we identified developmental changes in the genetic architecture of social-communication difficulties that enhance the understanding of ASD and schizophrenia-related dimensions. More generally, GSEM present a framework for modelling shared genetic aetiologies between phenotypes and can provide prior information with respect to patterns and continuity of trait-disorder overlap.

genetics

The genetic architecture of osteoarthritis: insights from UK Biobank

Osteoarthritis is a common complex disease with huge public health burden. Here we perform a genome-wide association study for osteoarthritis using data across 16.5 million variants from the UK Biobank resource. Following replication and meta-analysis in up to 30,727 cases and 297,191 controls, we report 9 new osteoarthritis loci, in all of which the most likely causal variant is non-coding. For three loci, we detect association with biologically-relevant radiographic endophenotypes, and in five signals we identify genes that are differentially expressed in degraded compared to intact articular cartilage from osteoarthritis patients. We establish causal effects for higher body mass index, but not for triglyceride levels or type 2 diabetes liability, on osteoarthritis.

genetics

Automating Mendelian randomization through machine learning to construct a putative causal map of the human phenome

A major application for genome-wide association studies (GWAS) has been the emerging field of causal inference using Mendelian randomization (MR), where the causal effect between a pair of traits can be estimated using only summary level data. MR depends on SNPs exhibiting vertical pleiotropy, where the SNP influences an outcome phenotype only through an exposure phenotype. Issues arise when this assumption is violated due to SNPs exhibiting horizontal pleiotropy. We demonstrate that across a range of pleiotropy models, instrument selection will be increasingly liable to selecting invalid instruments as GWAS sample sizes continue to grow. Methods have been developed in an attempt to protect MR from different patterns of horizontal pleiotropy, and here we have designed a mixture-of-experts machine learning framework (MR-MoE 1.0) that predicts the most appropriate model to use for any specific causal analysis, improving on both power and false discovery rates. Using the approach, we systematically estimated the causal effects amongst 2407 phenotypes. Almost 90% of causal estimates indicated some level of horizontal pleiotropy. The causal estimates are organised into a publicly available graph database (http://eve.mrbase.org), and we use it here to highlight the numerous challenges that remain in automated causal inference.

epidemiology

The Impact of Education on Myopia: A bidirectional Mendelian randomisation analysis in UK Biobank

Myopia, or short-sightedness, is one of the leading causes of visual disability in the World. The prevalence of myopia has risen steadily over recent decades, reaching epidemic levels in Southeast Asia. Observational studies have reported associations between educational attainment and myopia. Whether education causes myopia, myopic children are more intelligent, or another factor, like higher socioeconomic status, causes both is unclear since observational studies are prone to confounding and randomised trials of education are unethical. Using bidirectional Mendelian Randomisation, a form of instrumental variable (IV) analysis free from confounding, we show that every additional year in education leads to an increase in myopic refractive error, but that myopia does not lead to higher educational attainment. Our results suggest that current educational methods contribute to the global burden of myopia, and argue that educational policies and practices should take account of this to reduce future visual disability in the population.

epidemiology

Imprinted loci may be more widespread in humans than previously appreciated and enable limited assignment of parental allelic transmissions in unrelated individuals

Genomic imprinting is an epigenetic mechanism leading to parent-of-origin dependent gene expression. So far, the precise number of imprinted genes in humans is uncertain. In this study, we leveraged genome-wide DNA methylation in whole blood measured longitudinally at 3 time points (birth, childhood and adolescence) and GWAS data in 740 Mother-Child duos from the Avon Longitudinal Study of Parents and Children (ALSPAC) to systematically identify imprinted loci. We reasoned that cis-meQTLs at genomic regions that were imprinted would show strong evidence of parent-of-origin associations with DNA methylation, enabling the detection of imprinted regions. Using this approach, we identified genome-wide significant cis-meQTLs that exhibited parent-of-origin effects (POEs) at 35 novel and 50 known imprinted regions (10-10< P <10-300). Among the novel loci, we observed signals near genes implicated in cardiovascular disease (PCSK9), and Alzheimers disease (CR1), amongst others. Most of the significant regions exhibited imprinting patterns consistent with uniparental expression, with the exception of twelve loci (including the IGF2, IGF1R, and IGF2R genes), where we observed a bipolar-dominance pattern. POEs were remarkably consistent across time points and were so strong at some loci that methylation levels enabled good discrimination of parental transmissions at these and surrounding genomic regions. The implication is that parental allelic transmissions could be modelled at many imprinted (and linked) loci and hence POEs detected in GWAS of unrelated individuals given a combination of genetic and methylation data. Our results indicate that modelling POEs on DNA methylation is effective to identify loci that may be affected by imprinting.

genomics

Improving the accuracy of two-sample summary data Mendelian randomization: moving beyond the NOME assumption

BackgroundTwo-sample summary data Mendelian randomization (MR) incorporating multiple genetic variants within a meta-analysis framework is a popular technique for assessing causality in epidemiology. If all genetic variants satisfy the instrumental variable (IV) and necessary modelling assumptions, then their individual ratio estimates of causal effect should be homogeneous. Observed heterogeneity signals that one or more of these assumptions could have been violated.\n\nMethodsCausal estimation and heterogeneity assessment in MR requires an approximation for the variance, or equivalently the inverse-variance weight, of each ratio estimate. We show that the most popular 1st order weights can lead to an inflation in the chances of detecting heterogeneity when in fact it is not present. Conversely, ostensibly more accurate 2nd order weights can dramatically increase the chances of failing to detect heterogeneity, when it is truly present. We derive modified weights to mitigate both of these adverse effects.\n\nResultsUsing Monte Carlo simulations, we show that the modified weights outperform 1st and 2nd order weights in terms of heterogeneity quantification. Modified weights are also shown to remove the phenomenon of regression dilution bias in MR estimates obtained from weak instruments, unlike those obtained using 1st and 2nd order weights. However, with small numbers of weak instruments, this comes at the cost of a reduction in estimate precision and power to detect a causal effect compared to 1st order weighting. Moreover, 1st order weights always furnish unbiased estimates and preserve the type I error rate under the causal null. We illustrate the utility of the new method using data from a recent two-sample summary data MR analysis to assess the causal role of systolic blood pressure on coronary heart disease risk.\n\nConclusionsWe propose the use of modified weights within two-sample summary data MR studies for accurately quantifying heterogeneity and detecting outliers in the presence of weak instruments. Modified weights also have an important role to play in terms of causal estimation (in tandem with 1st order weights) but further research is required to understand their strengths and weaknesses in specific settings.

epidemiology

Investigating the role of insulin in increased adiposity: Bi-directional Mendelian randomization study

Insulin may serve as a key causal agent which regulates fat accumulation in the body. Here we assessed the causal relationship between fasting insulin and adiposity using publicly-available results from two large-scale genome-wide association studies for body mass index and fasting insulin levels in a two-sample, bidirectional Mendelian Randomized approach. This approach is only valid on the condition that the two instruments are independent of one another. In analysis excluding overlapping loci, there was an increase of 0.20 (0.17, 0.23) log pmol/L fasting insulin per SD increase in BMI (P= 2.80 x 10-36), while there was a null effect of fasting insulin on BMI, with a 0.01 (-0.39, 0.38) SD decrease in BMI per log pmol/L increase in fasting insulin (P= 0.98). Furthermore, a high degree of heterogeneity in the causal estimates was obtained from the insulin-related variants, which may be attributed to varying mechanisms of action of the insulin-associated variants. Results were largely consistent when an Egger regression technique and weighted median and mode estimators were applied. Findings suggest that the positive correlation between adiposity and fasting insulin levels are at least in part explained by the causal effect of adiposity on increasing insulin, rather than vice versa.

epidemiology