bioRxiv ScienceSearch

Biology subjects

Gusev, A.

Publications and source records attributed to Gusev, A..

10 recordsLinked to original sources

Large-scale transcriptome-wide association study identifies new prostate cancer risk regions

Although genome-wide association studies (GWAS) for prostate cancer (PrCa) have identified more than 100 risk regions, most of the risk genes at these regions remain largely unknown. Here, we integrate the largest PrCa GWAS (N=142,392) with gene expression measured in 45 tissues (N=4,458), including normal and tumor prostate, to perform a multi-tissue transcriptomewide association study (TWAS) for PrCa. We identify 235 genes at 87 independent 1Mb regions associated with PrCa risk, 9 of which are regions with no genome-wide significant SNP within 2Mb. 24 genes are significant in TWAS only for alternative splicing models in prostate tumor thus supporting the hypothesis of splicing driving risk for continued oncogenesis. Finally, we use a Bayesian probabilistic approach to estimate credible sets of genes containing the causal gene at pre-defined level; this reduced the list of 235 associations to 120 genes in the 90% credible set. Overall, our findings highlight the power of integrating expression with PrCa GWAS to identify novel risk loci and prioritize putative causal genes at known risk loci.

genetics

Multi-Tissue Transcriptome-Wide Association Studies Identify 21 Novel Candidate Susceptibility Genes for High Grade Serous Epithelial Ovarian Cancer

Genome-wide association studies (GWASs) have identified about 30 different susceptibility loci associated with high grade serous ovarian cancer (HGSOC) risk. We sought to identify potential susceptibility genes by integrating the risk variants at these regions with genetic variants impacting gene expression and splicing of nearby genes. We compiled gene expression and genotyping data from 2,169 samples for 6 different HGSOC-relevant tissue types. We integrated these data with GWAS data from 13,037 HGSOC cases and 40,941 controls, and performed a transcriptome-wide association study (TWAS) across >70,000 significantly heritable gene/exon features. We identified 24 transcriptome-wide significant associations for 14 unique genes, plus 90 significant exon-level associations in 20 unique genes. We implicated multiple novel genes at risk loci, e.g. LRRC46 at 19q21.32 (TWAS P=1x10-9) and a PRC1 splicing event (TWAS P=9x10-8) which was splice-variant specific and exhibited no eQTL signal. Functional analyses in HGSOC cell lines found evidence of essentiality for GOSR2, INTS1, KANSL1 and PRC1; with the latter gene showing levels of essentiality comparable to that of MYC. Overall, gene expression and splicing events explained 41% of SNP-heritability for HGSOC (s.e. 11%, P=2.5x10-4), implicated at least one target gene for 6/13 distinct genome-wide significant regions and revealed 2 known and 26 novel candidate susceptibility genes for HGSOC.\n\nSTATEMENT OF SIGNIFICANCEFor many ovarian cancer risk regions, the target genes regulated by germline genetic variants are unknown. Using expression data from >2,100 individuals, this study identified novel associations of genes and splicing variants with ovarian cancer risk; with transcriptional variation now explaining over one-third of the SNP-heritability for this disease.

genomics

Probabilistic fine-mapping of transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) using predicted expression have identified thousands of genes whose locally-regulated expression is associated to complex traits and diseases. In this work, we show that linkage disequilibrium (LD) among SNPs induce significant gene-trait associations at non-causal genes as a function of the overlap between eQTL weights used in expression prediction. We introduce a probabilistic framework that models the induced correlation among TWAS signals to assign a probability for every gene in the risk region to explain the observed association signal while controlling for pleiotropic SNP effects and unmeasured causal expression. Importantly, our approach remains accurate when expression data for causal genes are not available in the causal tissue by leveraging expression prediction from other tissues. Our approach yields credible-sets of genes containing the causal gene at a nominal confidence level (e.g., 90%) that can be used to prioritize and select genes for functional assays. We illustrate our approach using an integrative analysis of lipids traits where our approach prioritizes genes with strong evidence for causality.

genetics

Assessing the genetic effect mediated through gene expression from summary eQTL and GWAS data

Integrating genome-wide association (GWAS) and expression quantitative trait locus (eQTL) data into transcriptome-wide association studies (TWAS) based on predicted expression can boost power to detect novel disease loci or pinpoint the susceptibility gene at a known disease locus. However, it is often the case that multiple eQTL genes colocalize at disease loci, making the identification of the true susceptibility gene challenging, due to confounding through linkage disequilibrium (LD). To distinguish between true susceptibility genes (where the genetic effect on phenotype is mediated through expression) and colocalization due to LD, we examine an extension of the Mendelian Randomization Egger regression method that allows for LD while only requiring summary association data for both GWAS and eQTL. We derive the standard TWAS approach in the context of Mendelian Randomization and show in simulations that the standard TWAS does not control Type I error for causal gene identification when eQTLs have pleiotropic or LD-confounded effects on disease. In contrast, LD Aware MR-Egger regression can control Type I error in this case while attaining similar power as other methods in situations where these provide valid tests. However, when the direct effects of genetic variants on traits are correlated with the eQTL associations, all of the methods we examined including LD Aware MR-Egger regression can have inflated Type I error. We illustrate these methods by integrating gene expression within a recent large-scale breast cancer GWAS to provide guidance on susceptibility gene identification.

epidemiology

Detecting genome-wide directional effects of transcription factor binding on polygenic disease risk

Biological interpretation of GWAS data frequently involves analyzing unsigned genomic annotations comprising SNPs involved in a biological process and assessing enrichment for disease signal. However, it is often possible to generate signed annotations quantifying whether each SNP allele promotes or hinders a biological process, e.g., binding of a transcription factor (TF). Directional effects of such annotations on disease risk enable stronger statements about causal mechanisms of disease than enrichments of corresponding unsigned annotations. Here we introduce a new method, signed LD profile regression, for detecting such directional effects using GWAS summary statistics, and we apply the method using 382 signed annotations reflecting predicted TF binding. We show via theory and simulations that our method is well-powered and is well-calibrated even when TF binding sites co-localize with other enriched regulatory elements, which can confound unsigned enrichment methods. We further validate our method by showing that it recovers known transcriptional regulators when applied to molecular QTL in blood. We then apply our method to eQTL in 48 GTEx tissues, identifying 651 distinct TF-tissue expression associations at per-tissue FDR < 5%, including 30 associations with robust evidence of tissue specificity. Finally, we apply our method to 46 diseases and complex traits (average N = 289,617) and identify 77 annotation-trait associations at per-trait FDR < 5% representing 12 independent TF-trait associations, and we conduct gene-set enrichment analyses to characterize the underlying transcriptional programs. Our results implicate new causal disease genes (including causal genes at known GWAS loci), and in some cases suggest a detailed mechanism for a causal genes effect on disease. Our method provides a new way to leverage functional data to draw inferences about disease etiology.

genetics

Leveraging molecular QTL to understand the genetic architecture of diseases and complex traits

There is increasing evidence that many GWAS risk loci are molecular QTL for gene ex-pression (eQTL), histone modification (hQTL), splicing (sQTL), and/or DNA methylation (meQTL). Here, we introduce a new set of functional annotations based on causal posterior prob-abilities (CPP) of fine-mapped molecular cis-QTL, using data from the GTEx and BLUEPRINT consortia. We show that these annotations are very strongly enriched for disease heritability across 41 independent diseases and complex traits (average N = 320K): 5.84x for GTEx eQTL, and 5.44x for eQTL, 4.27-4.28x for hQTL (H3K27ac and H3K4me1), 3.61x for sQTL and 2.81x for meQTL in BLUEPRINT (all P [&le;] 1.39e-10), far higher than enrichments obtained using stan-dard functional annotations that include all significant molecular cis-QTL (1.17-1.80x). eQTL annotations that were obtained by meta-analyzing all 44 GTEx tissues generally performed best, but tissue-specific blood eQTL annotations produced stronger enrichments for autoimmune dis-eases and blood cell traits and tissue-specific brain eQTL annotations produced stronger enrich-ments for brain-related diseases and traits, despite high cis-genetic correlations of eQTL effect sizes across tissues. Notably, eQTL annotations restricted to loss-of-function intolerant genes from ExAC were even more strongly enriched for disease heritability (17.09x; vs. 5.84x for all genes; P = 4.90e-17 for difference). All molecular QTL except sQTL remained significantly enriched for disease heritability in a joint analysis conditioned on each other and on a broad set of functional annotations from previous studies, implying that each of these annotations is uniquely informative for disease and complex trait architectures.

genetics

Evidence for evolutionary shifts in the fitness landscape of human complex traits

Selection and mutation shape genetic variation underlying human traits, but the specific evolutionary mechanisms driving complex trait variation are largely unknown. We developed a statistical method that uses polarized GWAS summary statistics from a single population to detect signals of mutational bias and selection. We found evidence for non-neutral signals on variation underlying several traits (BMI, schizophrenia, Crohns disease, educational attainment, and height). We then used simulations that incorporate simultaneous negative and positive selection to show that these signals are consistent with mutational bias and shifts in the fitness-phenotype relationship, but not stabilizing selection or mutational bias alone. We additionally replicate two of our top three signals (BMI and educational attainment) in an external cohort, and show that population stratification may have confounded GWAS summary statistics for height in the GIANT cohort. Our results provide a flexible and powerful framework for evolutionary analysis of complex phenotypes in humans and other species, and offer insights into the evolutionary mechanisms driving variation in human polygenic traits.\n\nImpact summaryMany traits are variable within human populations and are likely to have a substantial and complex genetic component. This implies that mutations that have a functional impact on complex human traits have arisen throughout our species evolutionary history. However, it remains unclear how processes such as natural selection may have acted to shape trait variation at the genetic and phenotypic level. Better understanding of the mechanisms driving trait variation could provide insights into our evolutionary past and help clarify why it has been so difficult to map the preponderance of causal variation for common heritable diseases.\n\nIn this study, we developed and applied methods for detecting signatures of mutation bias (i.e., the propensity of a new variant to be either trait-increasing or trait-decreasing) and natural selection acting on trait variation. We applied our approach to several heritable traits, and found evidence for both natural selection and mutation bias, including selection for decreased BMI and decreased risk for Crohns disease and schizophrenia.\n\nWhile our results are consistent with plausible evolutionary scenarios shaping a range of traits, it should be noted that the field of polygenic selection detection is still new, and current methods (including ours) rely on data from genome-wide association studies (GWAS). The data produced by these studies may be vulnerable to certain cryptic biases, especially population stratification, which could induce false selection signals. We therefore repeated our analyses for the top three hits in a cohort that should be less susceptible to this problem - we found that two of our top three signals replicated (BMI and educational attainment), while height did not. Our results highlight both the promise and pitfalls of polygenic selection detection approaches, and suggest a need for further work disentangling stratification from selection.

evolutionary biology

Estimating the proportion of disease heritability mediated by gene expression levels

Disease risk variants identified by GWAS are predominantly noncoding, suggesting that gene regulation plays an important role. eQTL studies in unaffected individuals are often used to link disease-associated variants with the genes they regulate, relying on the hypothesis that noncoding regulatory effects are mediated by steady-state expression levels. To test this hypothesis, we developed a method to estimate the proportion of disease heritability mediated by the cis-genetic component of assayed gene expression levels. The method, gene expression co-score regression (GECS regression), relies on the idea that, for a gene whose expression level affects a phenotype, SNPs with similar effects on the expression of that gene will have similar phenotypic effects. In order to distinguish directional effects mediated by gene expression from non-directional pleiotropic or tagging effects, GECS regression operates on pairs of cis SNPs in linkage equilibrium, regressing pairwise products of disease effect sizes on products of cis-eQTL effect sizes. We verified that GECS regression produces robust estimates of mediated effects in simulations. We applied the method to eQTL data in 44 tissues from the GTEx consortium (average NeQTL = 158 samples) in conjunction with GWAS summary statistics for 30 diseases and complex traits (average NGWAS = 88K) with low pairwise genetic correlation, estimating the proportion of SNP-heritability mediated by the cis-genetic component of assayed gene expression in the union of the 44 tissues. The mean estimate was 0.21 (s.e. = 0.01) across 30 traits, with a significantly positive estimate (p < 0.001) for every trait. Thus, assayed gene expression in bulk tissues mediates a statistically significant but modest proportion of disease heritability, motivating the development of additional assays to capture regulatory effects and the use of our method to estimate how much disease heritability they mediate.

genetics

Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types

Genetics can provide a systematic approach to discovering the tissues and cell types relevant for a complex disease or trait. Identifying these tissues and cell types is critical for following up on non-coding allelic function, developing ex-vivo models, and identifying therapeutic targets. Here, we analyze gene expression data from several sources, including the GTEx and PsychENCODE consortia, together with genome-wide association study (GWAS) summary statistics for 48 diseases and traits with an average sample size of 169,331, to identify disease-relevant tissues and cell types. We develop and apply an approach that uses stratified LD score regression to test whether disease heritability is enriched in regions surrounding genes with the highest specific expression in a given tissue. We detect tissue-specific enrichments at FDR < 5% for 34 diseases and traits across a broad range of tissues that recapitulate known biology. In our analysis of traits with observed central nervous system enrichment, we detect an enrichment of neurons over other brain cell types for several brain-related traits, enrichment of inhibitory over excitatory neurons for bipolar disorder but excitatory over inhibitory neurons for schizophrenia and body mass index, and enrichments in the cortex for schizophrenia and in the striatum for migraine. In our analysis of traits with observed immunological enrichment, we identify enrichments of T cells for asthma and eczema, B cells for primary biliary cirrhosis, and myeloid cells for Alzheimer's disease, which we validated with independent chromatin data. Our results demonstrate that our polygenic approach is a powerful way to leverage gene expression data for interpreting GWAS signal.

genetics

Linkage disequilibrium dependent architecture of human complex traits reveals action of negative selection

Recent work has hinted at the linkage disequilibrium (LD) dependent architecture of human complex traits, where SNPs with low levels of LD (LLD) have larger per-SNP heritability after conditioning on their minor allele frequency (MAF). However, this has not been formally assessed, quantified or biologically interpreted. Here, we analyzed summary statistics from 56 complex diseases and traits (average N = 101,401) by extending stratified LD score regression to continuous annotations. We determined that SNPs with low LLD have significantly larger per-SNP heritability. Roughly half of the LLD signal can be explained by functional annotations that are negatively correlated with LLD, such as DNase I hypersensitivity sites (DHS). The remaining signal is largely driven by our finding that common variants that are more recent tend to have lower LLD and to explain more heritability (P = 2.38 x 10-104); the youngest 20% of common SNPs explain 3.9x more heritability than the oldest 20%, consistent with the action of negative selection. We also inferred jointly significant effects of other LD-related annotations and confirmed via forward simulations that these annotations jointly predict deleterious effects. Our results are consistent with the action of negative selection on deleterious variants that affect complex traits, complementing efforts to learn about negative selection by analyzing much smaller rare variant data sets.

genetics