bioRxiv ScienceSearch

Biology subjects

Price, A. L.

Publications and source records attributed to Price, A. L..

13 recordsLinked to original sources

Genes with high network connectivity are enriched for disease heritability

Recent studies have highlighted the role of gene networks in disease biology. To formally assess this, we constructed a broad set of pathway, network, and pathway+network annotations and applied stratified LD score regression to 42 independent diseases and complex traits (average N=323K) to identify enriched annotations. First, we constructed annotations from 18,119 biological pathways, including 100kb windows around each gene. We identified 156 pathway-trait pairs whose disease enrichment was statistically significant (FDR < 5%) after conditioning on all genes and on annotations from the baseline-LD model, a stringent step that greatly reduced the number of pathways detected; most of the significant pathway-trait pairs were previously unreported. Next, for each of four published gene networks, we constructed probabilistic annotations based on network connectivity using closeness centrality, a measure of how close a gene is to other genes in the network. For each gene network, the network connectivity annotation was strongly significantly enriched. Surprisingly, the enrichments were fully explained by excess overlap between network annotations and regulatory annotations from the baseline-LD model, validating the informativeness of the baseline-LD model and emphasizing the importance of accounting for regulatory annotations in gene network analyses. Finally, for each of the 156 enriched pathway-trait pairs, for each of the four gene networks, we constructed pathway+network annotations by annotating genes with high network connectivity to the input pathway. For each gene network, these pathway+network annotations were strongly significantly enriched for the corresponding traits. Once again, the enrichments were largely explained by the baseline-LD model. In conclusion, gene network connectivity is highly informative for disease architectures, but the information in gene networks may be subsumed by regulatory annotations, such that accounting for known annotations is critical to robust inference of biological mechanisms.

genetics

Disease heritability enrichment of regulatory elements is concentrated in elements with ancient sequence age and conserved function across species

Regulatory elements, e.g. enhancers and promoters, have been widely reported to be enriched for disease and complex trait heritability. We investigated how this enrichment varies with the age of the underlying genome sequence, the conservation of regulatory function across species, and the target gene of the regulatory element. We estimated heritability enrichment by applying stratified LD score regression to summary statistics from 41 independent diseases and complex traits (average N =320K) and meta-analyzing results across traits. Enrichment of human enhancers and promoters was larger in elements with older sequence age, assessed via alignment with other species irrespective of conserved functionality: enhancer elements with ancient sequence age (older than the split between marsupial and placental mammals) were 8.8x enriched (vs. 2.5x for all enhancers; p = 3e-14), and promoter elements with ancient sequence age were 13.5x enriched (vs. 5.1x for all promoters; p = 5e-16). Enrichment of human enhancers and promoters was also larger in elements whose regulatory function was conserved across species, e.g. human enhancers that were enhancers in [&ge;]5 of 9 other mammals were 4.6x enriched (p = 5e-12 vs. all enhancers). Enrichment of human promoters was larger in promoters of loss-of-function intolerant genes: 12.0x enrichment (p = 8e-15 vs. all promoters). The mean value of several measures of negative selection within these genomic annotations mirrored all of these findings. Notably, the annotations with these excess heritability enrichments were jointly significant conditional on each other and on our baseline-LD model, which includes a broad set of coding, conserved, regulatory and LD-related annotations.

genetics

Polygenicity of complex traits is explained by negative selection

Complex traits and common disease are highly polygenic: thousands of common variants are causal, and their effect sizes are almost always small. Polygenicity could be explained by negative selection, which constrains common-variant effect sizes and may reshape their distribution across the genome. We refer to this phenomenon as flattening, as genetic signal is flattened relative to the underlying biology. We introduce a mathematical definition of polygenicity, the effective number of associated SNPs, and a robust statistical method to estimate it. This definition of polygenicity differs from the number of causal SNPs, a standard definition; it depends strongly on SNPs with large effects. In analyses of 33 complex traits (average N=361k), we determined that common variants are [~]4x more polygenic than low-frequency variants, consistent with pervasive flattening. Moreover, functionally important regions of the genome have increased polygenicity in proportion to their increased heritability, implying that heritability enrichment reflects differences in the number of associations rather than their magnitude (which is constrained by selection). We conclude that negative selection constrains the genetic signal of biologically important regions and genes, reshaping genetic architecture.

genetics

GBAT: a gene-based association method for robust trans-gene regulation detection

Identification of trans-eQTLs has been limited by a heavy multiple testing burden, read-mapping biases, and hidden confounders. To address these issues, we developed GBAT, a powerful gene-based method that allows robust detection of trans gene regulation. Using simulated and real data, we show that GBAT drastically increases detection of trans-gene regulation over standard trans-eQTL analyses.

genetics

Modeling functional enrichment improves polygenic prediction accuracy in UK Biobank and 23andMe data sets

Genetic variants in functional regions of the genome are enriched for complex trait heritability. Here, we introduce a new method for polygenic prediction, LDpred-funct, that leverages trait-specific functional priors to increase prediction accuracy. We fit priors using the recently developed baseline-LD model, which includes coding, conserved, regulatory and LD-related annotations. We analytically estimate posterior mean causal effect sizes and then use cross-validation to regularize these estimates, improving prediction accuracy for sparse architectures. LDpred-funct attained higher prediction accuracy than other polygenic prediction methods in simulations using real genotypes. We applied LDpred-funct to predict 21 highly heritable traits in the UK Biobank. We used association statistics from British-ancestry samples as training data (avg N=373K) and samples of other European ancestries as validation data (avg N=22K), to minimize confounding. LDpred-funct attained a +4.6% relative improvement in average prediction accuracy (avg prediction R2=0.144; highest R2=0.413 for height) compared to SBayesR (the best method that does not incorporate functional information). For height, meta-analyzing training data from UK Biobank and 23andMe cohorts (total N=1107K; higher heritability in UK Biobank cohort) increased prediction R2 to 0.431. Our results show that incorporating functional priors improves polygenic prediction accuracy, consistent with the functional architecture of complex traits.

genetics

Quantification of genetic components of population differentiation in UK Biobank traits reveals signals of polygenic selection

The genetic architecture of most human complex traits is highly polygenic, motivating efforts to detect polygenic selection involving a large number of loci. In contrast to previous work relying on top GWAS loci, we developed a method that uses genome-wide association statistics and linkage disequilibrium patterns to estimate the genome-wide genetic component of population differentiation of a complex trait along a continuous gradient, enabling powerful inference of polygenic selection. We analyzed 43 UK Biobank traits and focused on PC1 and North-South and East-West birth coordinates across 337K unrelated British-ancestry samples, for which our method produced close to unbiased estimates of genetic components of population differentiation and high power to detect polygenic selection in simulations across different trait architectures. For PC1, we identified signals of polygenic selection for height (74.5{+/-}16.7% of 9.3% total correlation with PC1 attributable to genome-wide genetic effects; P = 8.4x10-6) and red hair pigmentation (95.9{+/-}24.7% of total correlation with PC1 attributable to genome-wide genetic effects; P = 1.1x10-4); the bulk of the signal remained when removing genome-wide significant loci, even though red hair pigmentation includes loci of large effect. We also detected polygenic selection for height, systolic blood pressure, BMI and basal metabolic rate along North-South birth coordinate, and height and systolic blood pressure along East-West birth coordinate. Our method detects polygenic selection in modern human populations with very subtle population structure and elucidates the relative contributions of genetic and non-genetic components of trait population differences.

genetics

High-throughput inference of pairwise coalescence times identifies signals of selection and enriched disease heritability

Interest in reconstructing demographic histories has motivated the development of methods to estimate locus-specific pairwise coalescence times from whole-genome sequence data. We developed a new method, ASMC, that can estimate coalescence times using only SNP array data, and is 2-4 orders of magnitude faster than previous methods when sequencing data are available. We were thus able to apply ASMC to 113,851 phased British samples from the UK Biobank, aiming to detect recent positive selection by identifying loci with unusually high density of very recent coalescence times. We detected 12 genome-wide significant signals, including 6 loci with previous evidence of positive selection and 6 novel loci, consistent with coalescent simulations showing that our approach is well-powered to detect recent positive selection. We also applied ASMC to sequencing data from 498 Dutch individuals (Genome of the Netherlands data set) to detect background selection at deeper time scales. We observed highly significant correlations between average coalescence time inferred by ASMC and other measures of background selection. We investigated whether this signal translated into an enrichment in disease and complex trait heritability by analyzing summary association statistics from 20 independent diseases and complex traits (average N=86k) using stratified LD score regression. Our background selection annotation based on average coalescence time was strongly enriched for heritability (p = 7x10-153) in a joint analysis conditioned on a broad set of functional annotations (including other background selection annotations), meta-analyzed across traits; SNPs in the top 20% of our annotation were 3.8x enriched for heritability compared to the bottom 20%. These results underscore the widespread effects of background selection on disease and complex trait heritability.

genetics

Reconciling S-LDSC and LDAK functional enrichment estimates

Recent work has highlighted the importance of accounting for linkage disequilibrium (LD)-dependent genetic architectures in analyses of heritability, motivating the development of the baseline-LD model used by stratified LD score regression (S-LDSC) and the LDAK model. Although both models include LD-dependent effects, they produce very different estimates of functional enrichment (with larger estimates using the baseline-LD model), leading to different interpretations of the functional architecture of complex traits. Here, we perform formal model comparisons and empirical analyses to reconcile these findings. First, by performing model comparisons using a likelihood approach, we determined that the baseline-LD model attains likelihoods across 16 UK Biobank traits that are substantially higher than the LDAK model. Second, we determined that S-LDSC using a combined model (unlike methods that use the LDAK or baseline-LD models) produces robust enrichment estimates in simulations under both the LDAK and baseline-LD models, validating the combined model as a gold standard. Third, in analyses of 16 UK Biobank traits, we determined that enrichment estimates obtained by S-LDSC using the combined model were nearly identical to those obtained by S-LDSC using the baseline-LD model (concordance correlation coefficient{rho} c = 0.99), but were larger than those obtained using LDAK ({rho}c = 0.54). Notably, LDAK enrichment estimates were much higher for a non-default version of LDAK that models SNPs in perfect LD differently by assigning non-zero weights to all SNPs. Our results support the use of the baseline-LD model and confirm the existence of functional annotations that are highly enriched for complex trait heritability.

genetics

Distinguishing genetic correlation from causation across 52 diseases and complex traits

Mendelian randomization (MR) is widely used to identify causal relationships among heritable traits, but it can be confounded by genetic correlations reflecting shared etiology. We propose a model in which a latent causal variable mediates the genetic correlation between two traits. Under the latent causal variable (LCV) model, trait 1 is fully genetically causal for trait 2 if it is perfectly genetically correlated with the latent causal variable, implying that the entire genetic component of trait 1 is causal for trait 2; it is partially genetically causal for trait 2 if it has a high genetic correlation with the latent variable, implying that part of the genetic component of trait 1 is causal for trait 2. To quantify the degree of partial genetic causality, we define the genetic causality proportion (gcp). We fit this model using mixed fourth moments E([Formula]12) and E([Formula]12) of marginal effect sizes for each trait, exploiting the fact that if trait 1 is causal for trait 2 then SNPs affecting trait 1 (large [Formula]) will have correlated effects on trait 2 (large 12), but not vice versa. We performed simulations under a wide range of genetic architectures and determined that LCV, unlike state-of-the-art MR methods, produced well-calibrated false positive rates and reliable gcp estimates in the presence of genetic correlations and asymmetric genetic architectures; we also determined that LCV is well-powered to detect a causal effect. We applied LCV to GWAS summary statistics for 52 traits (average N=331k), identifying partially or fully genetically causal effects (1% FDR) for 59 pairs of traits, including 30 pairs of traits with high gcp estimates (g[c]p > 0.6). Results consistent with the published literature included genetically causal effects on myocardial infarction (MI) for LDL, triglycerides and BMI. Novel findings included a genetically causal effect of LDL on bone mineral density, consistent with clinical trials of statins in osteoporosis. These results demonstrate that it is possible to distinguish between genetic correlation and causation using genetic data.

genetics

Mixed model association for biobank-scale data sets

Biobank-based genome-wide association studies are enabling exciting insights in complex trait genetics, but much uncertainty remains over best practices for optimizing statistical power and computational efficiency in GWAS while controlling confounders. Here, we introduce a much faster version of our BOLT-LMM Bayesian mixed model association method-- capable of running analyses of the full UK Biobank cohort in a few days on a single compute node--and show that it produces highly powered, robust test statistics when run on all 459K European samples (retaining related individuals). When used to conduct a GWAS for height in UK Biobank, BOLT-LMM achieved power equivalent to linear regression on 650K samples--a 93% increase in effective sample size versus the common practice of analyzing unrelated British samples using linear regression (UK Biobank documentation; Bycroft et al. bioRxiv). Across a broader set of 23 highly heritable traits, the total number of independent GWAS loci detected increased from 5,839 to 10,759, an 84% increase. We recommend the use of BOLT-LMM (retaining related individuals) for biobank-scale analyses, and we have publicly released BOLT-LMM summary association statistics for the 23 traits analyzed as a resource for all researchers.

genetics

Quantification of frequency-dependent genetic architectures and action of negative selection in 25 UK Biobank traits

Understanding the role of rare variants is important in elucidating the genetic basis of human diseases and complex traits. It is widely believed that negative selection can cause rare variants to have larger per-allele effect sizes than common variants. Here, we develop a method to estimate the minor allele frequency (MAF) dependence of SNP effect sizes. We use a model in which per-allele effect sizes have variance proportional to [p(1-p)], where p is the MAF and negative values of imply larger effect sizes for rare variants. We estimate by maximizing its profile likelihood in a linear mixed model framework using imputed genotypes, including rare variants (MAF >0.07%). We applied this method to 25 UK Biobank diseases and complex traits (N = 113,851). All traits produced negative estimates with 20 significantly negative, implying larger rare variant effect sizes. The inferred best-fit distribution of true values across traits had mean -0.38 (s.e. 0.02) and standard deviation 0.08 (s.e. 0.03), with statistically significant heterogeneity across traits (P = 0.0014). Despite larger rare variant effect sizes, we show that for most traits analyzed, rare variants (MAF <1%) explain less than 10% of total SNP-heritability. Using evolutionary modeling and forward simulations, we validated the model of MAF-dependent trait effects and estimated the level of coupling between fitness effects and trait effects. Based on this analysis an average genome-wide negative selection coefficient on the order of 10-4 or stronger is necessary to explain the values that we inferred.

genetics

Estimating the proportion of disease heritability mediated by gene expression levels

Disease risk variants identified by GWAS are predominantly noncoding, suggesting that gene regulation plays an important role. eQTL studies in unaffected individuals are often used to link disease-associated variants with the genes they regulate, relying on the hypothesis that noncoding regulatory effects are mediated by steady-state expression levels. To test this hypothesis, we developed a method to estimate the proportion of disease heritability mediated by the cis-genetic component of assayed gene expression levels. The method, gene expression co-score regression (GECS regression), relies on the idea that, for a gene whose expression level affects a phenotype, SNPs with similar effects on the expression of that gene will have similar phenotypic effects. In order to distinguish directional effects mediated by gene expression from non-directional pleiotropic or tagging effects, GECS regression operates on pairs of cis SNPs in linkage equilibrium, regressing pairwise products of disease effect sizes on products of cis-eQTL effect sizes. We verified that GECS regression produces robust estimates of mediated effects in simulations. We applied the method to eQTL data in 44 tissues from the GTEx consortium (average NeQTL = 158 samples) in conjunction with GWAS summary statistics for 30 diseases and complex traits (average NGWAS = 88K) with low pairwise genetic correlation, estimating the proportion of SNP-heritability mediated by the cis-genetic component of assayed gene expression in the union of the 44 tissues. The mean estimate was 0.21 (s.e. = 0.01) across 30 traits, with a significantly positive estimate (p < 0.001) for every trait. Thus, assayed gene expression in bulk tissues mediates a statistically significant but modest proportion of disease heritability, motivating the development of additional assays to capture regulatory effects and the use of our method to estimate how much disease heritability they mediate.

genetics

Linkage disequilibrium dependent architecture of human complex traits reveals action of negative selection

Recent work has hinted at the linkage disequilibrium (LD) dependent architecture of human complex traits, where SNPs with low levels of LD (LLD) have larger per-SNP heritability after conditioning on their minor allele frequency (MAF). However, this has not been formally assessed, quantified or biologically interpreted. Here, we analyzed summary statistics from 56 complex diseases and traits (average N = 101,401) by extending stratified LD score regression to continuous annotations. We determined that SNPs with low LLD have significantly larger per-SNP heritability. Roughly half of the LLD signal can be explained by functional annotations that are negatively correlated with LLD, such as DNase I hypersensitivity sites (DHS). The remaining signal is largely driven by our finding that common variants that are more recent tend to have lower LLD and to explain more heritability (P = 2.38 x 10-104); the youngest 20% of common SNPs explain 3.9x more heritability than the oldest 20%, consistent with the action of negative selection. We also inferred jointly significant effects of other LD-related annotations and confirmed via forward simulations that these annotations jointly predict deleterious effects. Our results are consistent with the action of negative selection on deleterious variants that affect complex traits, complementing efforts to learn about negative selection by analyzing much smaller rare variant data sets.

genetics