bioRxiv ScienceSearch

Biology subjects

Kutalik, Z.

Publications and source records attributed to Kutalik, Z..

7 recordsLinked to original sources

Mendelian Randomization integrating GWAS and eQTL data reveals genetic determinants of complex and clinical traits

Genome-wide association studies (GWAS) identified thousands of variants associated with complex traits, but their biological interpretation often remains unclear. Most of these variants overlap with expression QTLs (eQTLs), indicating their potential involvement in the regulation of gene expression.\n\nHere, we propose an advanced transcriptome-wide summary statistics-based Mendelian Randomization approach (called TWMR) that uses multiple SNPs jointly as instruments and multiple gene expression traits as exposures, simultaneously.\n\nWhen applied to 43 human phenotypes it uncovered 2,369 genes whose blood expression is putatively associated with at least one phenotype resulting in 3,913 gene-trait associations; of note, 36% of them had no genome-wide significant SNP nearby in previous GWAS analysis. Using independent association summary statistics (UKBiobank), we confirmed that the majority of these loci were missed by conventional GWAS due to power issues. Noteworthy among these novel links is educational attainment-associated BSCL2, known to carry mutations leading to a mendelian form of encephalopathy. We similarly unraveled novel pleiotropic causal effects suggestive of mechanistic connections, e.g. the shared genetic effects of GSDMB in rheumatoid arthritis, ulcerative colitis and Crohns disease.\n\nOur advanced Mendelian Randomization unlocks hidden value from published GWAS through higher power in detecting associations. It better accounts for pleiotropy and unravels new biological mechanisms underlying complex and clinical traits.

genetics

Genomic underpinnings of lifespan allow prediction and reveal basis in modern risks

We use a multi-stage genome-wide association of 1 million parental lifespans of genotyped subjects and data on mortality risk factors to validate previously unreplicated findings near CDKN2B-AS1, ATXN2/BRAP, FURIN/FES, ZW10, PSORS1C3, and 13q21.31, and identify and replicate novel findings near GADD45G, KCNK3, LDLR, POM121C, ZC3HC1, and ABO. We also validate previous findings near 5q33.3/EBF1 and FOXO3, whilst finding contradictory evidence at other loci. Gene set and tissue-specific analyses show that expression in foetal brain cells and adult dorsolateral prefrontal cortex is enriched for lifespan variation, as are gene pathways involving lipid proteins and homeostasis, vesicle-mediated transport, and synaptic function. Individual genetic variants that increase dementia, cardiovascular disease, and lung cancer -but not other cancers-explain the most variance, possibly reflecting modern susceptibilities, whilst cancer may act through many rare variants, or the environment. Resultant polygenic scores predict a mean lifespan difference of around five years of life across the deciles.

genomics

Genetic studies of accelerometer-based sleep measures in 85,670 individuals yield new insights into human sleep behaviour

Sleep is an essential human function but its regulation is poorly understood. Identifying genetic variants associated with quality, quantity and timing of sleep will provide biological insights into the regulation of sleep and potential links with disease. Using accelerometer data from 85,670 individuals in the UK Biobank, we performed a genome-wide association study of 8 accelerometer-derived sleep traits, 5 of which are not accessible through self-report alone. We identified 47 genetic associations across the sleep traits (P<5x10-8) and replicated our findings in 5,819 individuals from 3 independent studies. These included 26 novel associations for sleep quality and 10 for nocturnal sleep duration. The majority of newly identified variants were associated with a single sleep trait, except for variants previously associated with restless legs syndrome that were associated with multiple sleep traits. Of the new associated and replicated sleep duration loci, we were able to fine-map a missense variant (p.Tyr727Cys) in PDE11A, a dual-specificity 3,5-cyclic nucleotide phosphodiesterase expressed in the hippocampus, as the likely causal variant. As a group, sleep quality loci were enriched for serotonin processing genes and all sleep traits were enriched for cerebellar-expressed genes. These findings provide new biological insights into sleep characteristics.

genetics

Cross-species functional modules link proteostasis to human normal aging

The evolutionarily conserved nature of the few well-known anti-aging interventions that affect lifespan, such as caloric restriction, suggests that aging-related research in model organisms is directly relevant to human aging. Since human lifespan is a complex trait, a systems-level approach will contribute to a more comprehensive understanding of the underlying aging landscape. Here, we integrate evolutionary and functional information of normal aging across human and model organisms at three levels: gene-level, process-level, and network-level. We identify evolutionarily conserved modules of normal aging across diverse taxa, and importantly, we show that proteostasis involvement is conserved in healthy aging. Additionally, we find that mechanisms related to protein quality control network are enriched in 22 age-related genome-wide association studies (GWAS) and are associated to caloric restriction. These results demonstrate that a systems-level approach, combined with evolutionary conservation, allows the detection of candidate aging genes and pathways relevant to human normal aging.\n\nHighlightsO_LINormal aging is evolutionarily conserved at the module level.\nC_LIO_LICore pathways in healthy aging are related to mechanisms of protein quality network\nC_LIO_LIThe evolutionarily conserved pathways of healthy aging react to caloric restriction.\nC_LIO_LIOur integrative approach identifies evolutionarily conserved functional modules and showed enrichment in several age-related GWAS studies.\nC_LI

systems biology

Open Community Challenge Reveals Molecular Network Modules with Key Roles in Diseases

Identification of modules in molecular networks is at the core of many current analysis methods in biomedical research. However, how well different approaches identify disease-relevant modules in different types of gene and protein networks remains poorly understood. We launched the "Disease Module Identification DREAM Challenge", an open competition to comprehensively assess module identification methods across diverse protein-protein interaction, signaling, gene co-expression, homology, and cancer-gene networks. Predicted network modules were tested for association with complex traits and diseases using a unique collection of 180 genome-wide association studies (GWAS). Our critical assessment of 75 contributed module identification methods reveals novel top-performing algorithms, which recover complementary trait-associated modules. We find that most of these modules correspond to core disease-relevant pathways, which often comprise therapeutic targets and correctly prioritize candidate disease genes. This community challenge establishes benchmarks, tools and guidelines for molecular network analysis to study human disease biology (https://synapse.org/modulechallenge).

bioinformatics

Evaluation and application of summary statistic imputation to discover new height-associated loci

AbstractAs most of the heritability of complex traits is attributed to common and low frequency genetic variants, imputing them by combining genotyping chips and large sequenced reference panels is the most cost-effective approach to discover the genetic basis of these traits. Association summary statistics from genome-wide meta-analyses are available for hundreds of traits. Updating these to ever-increasing reference panels is very cumbersome as it requires reimputation of the genetic data, rerunning the association scan, and meta-analysing the results. A much more efficient method is to directly impute the summary statistics, termed as summary statistics imputation. Its performance relative to genotype imputation and practical utility has not yet been fully investigated. To this end, we compared the two approaches on real (genotyped and imputed) data from 120K samples from the UK Biobank and show that, while genotype imputation boasts a 2- to 5-fold lower root-mean-square error, summary statistics imputation better distinguishes true associations from null ones: We observed the largest differences in power for variants with low minor allele frequency and low imputation quality. For fixed false positive rates of 0.001, 0.01, 0.05, using summary statistics imputation yielded an increase in statistical power by 15, 10 and 3%, respectively. To test its capacity to discover novel associations, we applied summary statistics imputation to the GIANT height meta-analysis summary statistics covering HapMap variants, and identified 34 novel loci, 19 of which replicated using data in the UK Biobank. Additionally, we successfully replicated 55 out of the 111 variants published in an exome chip study. Our study demonstrates that summary statistics imputation is a very efficient and cost-effective way to identify and fine-map trait-associated loci. Moreover, the ability to impute summary statistics is important for follow-up analyses, such as Mendelian randomisation or LD-score regression.\n\nAuthor summaryGenome-wide association studies (GWASs) quantify the effect of genetic variants and traits, such as height. Such estimates are called association summary statistics and are typically publicly shared through publication. Typically, GWASs are carried out by genotyping ~ 500'000 SNVs for each individual which are then combined with sequenced reference panels to infer untyped SNVs in each individuals genome. This process of genotype imputation is resource intensive and can therefore be a limitation when combining many GWASs. An alternative approach is to bypass the use of individual data and directly impute summary statistics. In our work we compare the performance of summary statistics imputation to genotype imputation. Although we observe a 2- to 5-fold lower RMSE for genotype imputation compared to summary statistics imputation, summary statistics imputation better distinguishes true associations from null results. Furthermore, we demonstrate the potential of summary statistics imputation by presenting 34 novel height-associated loci, 19 of which were confirmed in UK Biobank. Our study demonstrates that given current reference panels, summary statistics imputation is a very efficient and cost-effective way to identify common or low-frequency trait-associated loci.

genomics

Improved imputation of summary statistics for realistic settings

MotivationSummary statistics imputation can be used to infer association summary statistics of an already conducted, genotype-based meta-analysis to higher ge-nomic resolution. This is typically needed when genotype imputation is not feasible for some cohorts. Oftentimes, cohorts of such a meta-analysis are variable in terms of (country of) origin or ancestry. This violates the assumption of current methods that an external LD matrix and the covariance of the Z-statistics are identical.\n\nResultsTo address this issue, we present variance matching, an extention to the existing summary statistics imputation method, which manipulates the LD matrix needed for summary statistics imputation. Based on simulations using real data we find that accounting for ancestry admixture yields noticeable improvement only when the total reference panel size is > 1000. We show that for population specific variants this effect is more pronounced with increasing FST.

genomics