bioRxiv ScienceSearch

Biology subjects

Laakso, M.

Publications and source records attributed to Laakso, M..

9 recordsLinked to original sources

Reverse GWAS: Using Genetics to Identify and Model Phenotypic Subtypes

Recent and classical work has revealed biologically and medically significant subtypes in complex diseases and traits. However, relevant subtypes are often unknown, unmeasured, or actively debated, making automatic statistical approaches to subtype definition particularly valuable. We propose reverse GWAS (RGWAS) to identify and validate subtypes using genetics and multiple traits: while GWAS seeks the genetic basis of a given trait, RGWAS seeks to define trait subtypes with distinct genetic bases. Unlike existing approaches relying on off-the-shelf clustering methods, RGWAS uses a bespoke decomposition, MFMR, to model covariates, binary traits, and population structure. We use extensive simulations to show these features can be crucial for power and calibration. We validate RGWAS in practice by recovering known stress subtypes in major depressive disorder. We then show the utility of RGWAS by identifying three novel subtypes of metabolic traits. We biologically validate these metabolic subtypes with SNP-level tests and a novel polygenic test: the former recover known metabolic GxE SNPs; the latter suggests genetic heterogeneity may explain substantial missing heritability. Crucially, statins, which are widely prescribed and theorized to increase diabetes risk, have opposing effects on blood glucose across metabolic subtypes, suggesting potential have potential translational value.\n\nAuthor summaryComplex diseases depend on interactions between many known and unknown genetic and environmental factors. However, most studies aggregate these strata and test for associations on average across samples, though biological factors and medical interventions can have dramatically different effects on different people. Further, more-sophisticated models are often infeasible because relevant sources of heterogeneity are not generally known a priori. We introduce Reverse GWAS to simultaneously split samples into homogeneoues subtypes and to learn differences in genetic or treatment effects between subtypes. Unlike existing approaches to computational subtype identification using high-dimensional trait data, RGWAS accounts for covariates, binary disease traits and, especially, population structure; these features are each invaluable in extensive simulations. We validate RGWAS by recovering known genetic subtypes of major depression. We demonstrate RGWAS is practically useful in a metabolic study, finding three novel subtypes with both SNP- and polygenic-level heterogeneity. Importantly, RGWAS can uncover differential treatment response: for example, we show that statin, a common drug and potential type 2 diabetes risk factor, may have opposing subtype-specific effects on blood glucose.

genetics

Clustering of Type 2 Diabetes Genetic Loci by Multi-Trait Associations Identifies Disease Mechanisms and Subtypes

BackgroundType 2 diabetes (T2D) is a heterogeneous disease for which 1) disease-causing pathways are incompletely understood and 2) sub-classification may improve patient management. Unlike other biomarkers, germline genetic markers do not change with disease progression or treatment. In this paper we test whether a germline genetic approach informed by physiology can be used to deconstruct T2D heterogeneity. First, we aimed to categorize genetic loci into groups representing likely disease mechanistic pathways. Second, we asked whether the novel clusters of genetic loci we identified have any broad clinical consequence, as assessed in four independent cohorts of individuals with T2D.\n\nMethods and FindingsIn an effort to identify mechanistic pathways driven by established T2D genetic loci, we applied Bayesian nonnegative matrix factorization clustering to genome-wide association results for 94 independent T2D genetic loci and 47 diabetes-related traits. We identified five robust clusters of T2D loci and traits, each with distinct tissue-specific enhancer enrichment based on analysis of epigenomic data from 28 cell types. Two clusters contained variant-trait associations indicative of reduced beta-cell function, differing from each other by high vs. low proinsulin levels. The three other clusters displayed features of insulin resistance: obesity-mediated (high BMI, waist circumference), \"lipodystrophy-like\" fat distribution (low BMI, adiponectin, HDL-cholesterol, and high triglycerides), and disrupted liver lipid metabolism (low triglycerides). Increased cluster GRSs were associated with distinct clinical outcomes, including increased blood pressure, coronary artery disease, and stroke risk. We evaluated the potential for clinical impact of these clusters in four studies containing participants with T2D (METSIM, N=487; Ashkenazi, N=509; Partners Biobank, N=2,065; UK Biobank N=14,813). Individuals with T2D in the top genetic risk score decile for each cluster reproducibly exhibited the predicted cluster-associated phenotypes, with ~30% of all participants assigned to just one cluster top decile.\n\nConclusionOur approach identifies salient T2D genetically anchored and physiologically informed pathways, and supports use of genetics to deconstruct T2D heterogeneity. Classification of patients by these genetic pathways may offer a step toward genetically informed T2D patient management.

genetics

Discovery of biomarkers for glycaemic deterioration before and after the onset of type 2 diabetes: an overview of the data from the epidemiological studies within the IMI DIRECT Consortium

Abstract/SummaryO_ST_ABSBackground and aimsC_ST_ABSUnderstanding the aetiology, clinical presentation and prognosis of type 2 diabetes (T2D) and optimizing its treatment might be facilitated by biomarkers that help predict a persons susceptibility to the risk factors that cause diabetes or its complications, or response to treatment. The IMI DIRECT (Diabetes Research on Patient Stratification) Study is a European Union (EU) Innovative Medicines Initiative (IMI) project that seeks to test these hypotheses in two recently established epidemiological cohorts. Here, we describe the characteristics of these cohorts at baseline and at the first main follow-up examination (18-months).\n\nMaterials and methodsFrom a sampling-frame of 24,682 European-ancestry adults in whom detailed health information was available, participants at varying risk of glycaemic deterioration were identified using a risk prediction algorithm and enrolled into a prospective cohort study (n=2127) undertaken at four study centres across Europe (Cohort 1: prediabetes). We also recruited people from clinical registries with recently diagnosed T2D (n=789) into a second cohort study (Cohort 2: diabetes). The two cohorts were studied in parallel with matched protocols. Endogenous insulin secretion and insulin sensitivity were modelled from frequently sampled 75g oral glucose tolerance (OGTT) in Cohort 1 and with mixed-meal tolerance tests (MMTT) in Cohort 2. Additional metabolic biochemistry was determined using blood samples taken when fasted and during the tolerance tests. Body composition was assessed using MRI and lifestyle measures through self-report and objective methods.\n\nResultsUsing ADA-2011 glycaemic categories, 33% (n=693) of Cohort 1 (prediabetes) had normal glucose regulation (NGR), and 67% (n=1419) had impaired glucose regulation (IGR). 76% of the cohort was male, age=62(6.2) years; BMI=27.9(4.0) kg/m2; fasting glucose=5.7(0.6) mmol/l; 2-hr glucose=5.9(1.6) mmol/l [mean(SD)]. At follow-up, 18.6(1.4) months after baseline, fasting glucose=5.8(0.6) mmol/l; 2-hr OGTT glucose=6.1(1.7) mmol/l [mean(SD)]. In Cohort 2 (diabetes): 65% (n=508) were lifestyle treated (LS) and 35% (n=271) were lifestyle + metformin treated (LS+MET). 58% of the cohort was male, age=62(8.1) years; BMI=30.5(5.0) kg/m2; fasting glucose=7.2(1.4)mmol/l; 2-hr glucose=8.6(2.8) mmol/l [mean(SD)]. At follow-up, 18.2(0.6) months after baseline, fasting glucose=7.8(1.8) mmol/l; 2-hr MMTT glucose=9.5(3.3) mmol/l [mean(SD)].\n\nConclusionThe epidemiological IMI DIRECT cohorts are the most intensely characterised prospective studies of glycaemic deterioration to date. Data from these cohorts help illustrate the heterogeneous characteristics of people at risk of or with T2D, highlighting the rationale for biomarker stratification of the disease - the primary objective of the IMI DIRECT consortium.\n\nAbbreviations

epidemiology

Epigenome-wide association in adipose tissue from the METSIM cohort identifies novel loci and the involvement of adipocytes and macrophages in diabetes traits

Most epigenome-wide association studies to date have been conducted in blood. However, metabolic syndrome is mediated by a dysregulation of adiposity and therefore it is critical to study adipose tissue in order to understand the effects of this syndrome on epigenomes. To determine if natural variation in DNA methylation was associated with metabolic syndrome traits, we profiled global methylation levels in subcutaneous abdominal adipose tissue. We measured association between 32 clinical traits related to diabetes and obesity in 201 people from the Metabolic Syndrome In Men cohort. We performed epigenome-wide association studies between DNA methylation levels and traits, and identified associations for 13 clinical traits in 21 loci. We prioritized candidate genes in these loci using eQTL, and identified 18 high confidence candidate genes, including known and novel genes associated with diabetes and obesity traits. Using methylation deconvolution, we examined which cell types may be mediating the associations, and concluded that most of the loci we identified were specific to adipocytes. We determined whether the abundance of cell types varies with metabolic traits, and found that macrophages increased in abundance with the severity of metabolic syndrome traits. Finally, we developed a DNA methylation based biomarker to assess type II diabetes risk in adipose tissue. In conclusion, our results demonstrate that profiling DNA methylation in adipose tissue is a powerful tool for understanding the molecular effects of metabolic syndrome on adipose tissue, and can be used in conjunction with traditional genetic analyses to further characterize this disorder.

genetics

Proper Conditional Analysis in the Presence of Missing Data Identified Novel Independently Associated Low Frequency Variants in Nicotine Dependence Genes

Meta-analysis of genetic association studies increases sample size and the power for mapping complex traits. Existing methods are mostly developed for datasets without missing values. In practice, genotype imputation is not always effective, e.g. when targeted genotyping/sequencing assays are used or when the un-typed genetic variant is rare. Therefore, contributed summary statistics often contain missing values. Naive extensions of existing methods either replace missing summary statistics with 0 or discard studies with missing data. These approaches can bias genetic effect estimates and lead to seriously inflated type-I or II errors in conditional analysis, which is a critical tool for identifying independently associated variants.\n\nTo address this challenge and complement imputation methods, we developed a method to combine summary statistics across participating studies and consistently estimate joint effects, even when the contributed summary statistics contain large amount of missing values. Based on this estimator, we propose a score statistic we call PCBS (partial correlation based score statistic) for conditional analysis of single-variant and gene-level associations. Through extensive analysis of simulated and real data, we showed that the new method produces well-calibrated type-I errors and is substantially more powerful than existing approaches. We applied the proposed approach to analyze the CHRNA5-CHRNB4-CHRNA3 locus in a large-scale meta-analysis for cigarettes-per-day. Using the new method, we identified three novel variants, independent of known association signals, which were otherwise missed by alternative methods. Together, the phenotypic variance explained by these variants is .46%, improving that of previously reported associations by 17%. These findings illustrate the extent of locus allelic heterogeneity and can help pinpoint causal variants.\n\nAUTHOR SUMMARYIt is of great interest to estimate the joint and conditional effects of multiple correlated variants from large scale meta-analysis, in order to fine map causal variants and understand the genetic architecture for complex traits. The contributed summary statistics from participating studies in a meta-analysis often contain missing values, as the imputation methods are not often effective, especially when the underlying genetic variant is rare or the participating studies use targeted genotyping array that is not suitable for imputation. Existing meta-analysis methods do not properly handle missing data, and can incorrectly estimate correlations between score statistics. As a result, they can produce highly biased estimates of joint effects and highly inflated type-I errors for conditional analysis, which will in turn result in overestimated phenotypic variance explained and incorrect identification of causal variants. We systematically evaluated this bias and proposed a novel partial correlation based score statistic. The new statistic has valid type-I errors for conditional analysis and much higher power than the existing methods, even when the contributed summary statistics in the meta-analysis contain a large fraction of missing values. We expect this method to be highly useful in the sequencing age for complex trait genetics.

genetics

Association Analysis and Meta-Analysis of Multi-allelic Variants for Large Scale Sequence Data

MotivationThere is great interest to understand the impact of rare variants in human diseases using large sequence datasets. In deep sequences datasets of >10,000 samples, [~]10% of the variant sites are observed to be multi-allelic. Many of the multi-allelic variants have been shown to be functional and disease relevant. Proper analysis of multi-allelic variants is critical to the success of a sequencing study, but existing methods do not properly handle multi-allelic variants and can produce highly misleading association results.\n\nResultsWe propose novel methods to encode multi-allelic sites, conduct single variant and gene-level association analyses, and perform meta-analysis for multi-allelic variants. We evaluated these methods through extensive simulations and the study of a large meta-analysis of [~]18,000 samples on the cigarettes-per-day phenotype. We showed that our joint modeling approach provided an unbiased estimate of genetic effects, greatly improved the power of single variant association tests, and enhanced gene-level tests over existing approaches.\n\nAvailabilitySoftware packages implementing these methods are available at (https://github.com/zhanxw/rvtests http://genome.sph.umich.edu/wiki/RareMETAL).\n\nContactxiaowei.zhan@utsouthwestem.edu; dajiang.liu@psu.edu

bioinformatics

Quantifying the impact of rare and ultra-rare coding variation across the phenotypic spectrum

There is a limited understanding about the impact of rare protein truncating variants across multiple phenotypes. We explore the impact of this class of variants on 13 quantitative traits and 10 diseases using whole-exome sequencing data from 100,296 individuals. Protein truncating variants in genes intolerant to this class of mutations increased risk of autism, schizophrenia, bipolar disorder, intellectual disability, ADHD. In individuals without these disorders, there was an association with shorter height, lower education, increased hospitalization and reduced age. Gene sets implicated from GWAS did not show a significant protein truncating variants-burden beyond what captured by established Mendelian genes. In conclusion, we provide the most thorough investigation to date of the impact of rare deleterious coding variants on complex traits, suggesting widespread pleiotropic risk.\n\nMain abbreviations

genetics

Interactions between genetic variation and cellular environment in skeletal muscle gene expression

From whole organisms to individual cells, responses to environmental conditions are influenced by genetic makeup, where the effect of genetic variation on a trait depends on the environmental context. RNA-sequencing quantifies gene expression as a molecular trait, and is capable of capturing both genetic and environmental effects. In this study, we explore opportunities of using allele-specific expression (ASE) to discover cis acting genotype-environment interactions (GxE) - genetic effects on gene expression that depend on an environmental condition. Treating 17 common, clinical traits as approximations of the cellular environment of 267 skeletal muscle biopsies, we identify 10 candidate interaction quantitative trait loci (iQTLs) across 6 traits (12 unique gene-environment trait pairs; 10% FDR per trait) including sex, systolic blood pressure, and low-density lipoprotein cholesterol. Although using ASE is in principle a promising approach to detect GxE effects, replication of such signals can be challenging as validation requires harmonization of environmental traits across cohorts and a sufficient sampling of heterozygotes for a transcribed SNP. Comprehensive discovery and replication will require large human transcriptome datasets, or the integration of multiple transcribed SNPs, coupled with standardized clinical phenotyping.

genetics

Heterozygous RFX6 protein truncating variants cause Maturity-Onset Diabetes of the Young (MODY) with reduced penetrance

Finding new genetic causes of monogenic diabetes can help to understand development and function of the human pancreas. We aimed to find novel protein-truncating variants causing Maturity-Onset Diabetes of the Young (MODY), a subtype of monogenic diabetes. We used a combination of next-generation sequencing of MODY cases with unknown aetiology along with comparisons to the ExAC database to identify new MODY genes. In the discovery cohort of 36 European patients, we identified two probands with novel RFX6 heterozygous nonsense variants. RFX6 protein truncating variants were enriched in the MODY discovery cohort compared to the European control population within ExAC (odds ratio, OR=131, P=lxl0-4). We found similar results in non-Finnish European (n=348, OR=43, P=5xl0-5) and Finnish (n=80, OR=22, P=1xl0-6) replication cohorts. The overall meta-analysis OR was 34 (P=lxl0-16). RFX6 heterozygotes had reduced penetrance of diabetes compared to common HNF1A and HNF4A-MODY mutations (27%, 70% and 55% at 25 years of age, respectively). The hyperglycaemia resulted from beta-cell dysfunction and was associated with lower fasting and stimulated gastric inhibitory polypeptide (GIP) levels. Our study demonstrates that heterozygous RFX6 protein truncating variants are associated with MODY with reduced penetrance.

genetics