bioRxiv ScienceSearch

Biology subjects

Neale, B.

Publications and source records attributed to Neale, B..

13 recordsLinked to original sources

Comparative genetic architectures of schizophrenia in East Asian and European populations

Author summarySchizophrenia is a severe psychiatric disorder with a lifetime risk of about 1% world-wide. Most large schizophrenia genetic studies have studied people of primarily European ancestry, potentially missing important biological insights. Here we present a study of East Asian participants (22,778 schizophrenia cases and 35,362 controls), identifying 21 genome-wide significant schizophrenia associations in 19 genetic loci. Over the genome, the common genetic variants that confer risk for schizophrenia have highly similar effects in those of East Asian and European ancestry (rg=0.98), indicating for the first time that the genetic basis of schizophrenia and its biology are broadly shared across these world populations. A fixed-effect meta-analysis including individuals from East Asian and European ancestries revealed 208 genome-wide significant schizophrenia associations in 176 genetic loci (53 novel). Trans-ancestry fine-mapping more precisely isolated schizophrenia causal alleles in 70% of these loci. Despite consistent genetic effects across populations, polygenic risk models trained in one population have reduced performance in the other, highlighting the importance of including all major ancestral groups with sufficient sample size to ensure the findings have maximum relevance for all populations.

genetics

Signals of polygenic adaptation on height have been overestimated due to uncorrected population structure in genome-wide association studies

Genetic predictions of height differ among human populations and these differences are too large to be explained by genetic drift. This observation has been interpreted as evidence of polygenic adaptation. Differences across populations were detected using SNPs genome-wide significantly associated with height, and many studies also found that the signals grew stronger when large numbers of subsignificant SNPs were analyzed. This has led to excitement about the prospect of analyzing large fractions of the genome to detect subtle signals of selection and claims of polygenic adaptation for multiple traits. Polygenic adaptation studies of height have been based on SNP effect size measurements in the GIANT Consortium meta-analysis. Here we repeat the height analyses in the UK Biobank, a much more homogeneously designed study. Our results show that polygenic adaptation signals based on large numbers of SNPs below genome-wide significance are extremely sensitive to biases due to uncorrected population structure.

evolutionary biology

GWAS meta-analysis highlights the hypothalamic-pituitary-gonadal axis (HPG axis) in the genetic regulation of menstrual cycle length

The normal menstrual cycle requires a delicate interplay between the hypothalamus, pituitary, and ovary. Therefore, its length is an important indicator of female reproductive health. Menstrual cycle length has been shown to be partially controlled by genetic factors, especially in the follicle stimulating hormone beta-subunit (FSHB) locus. GWAS meta-analysis of menstrual cycle length in 44,871 women of European ancestry confirmed the previously observed association with the FSHB locus and identified four additional novel signals in, or near, the GNRH1, PGR, NR5A2 and INS-IGF2 genes. These findings confirm the role of the hypothalamic-pituitary-gonadal axis in the genetic regulation of menstrual cycle length, but also highlight potential novel local regulatory mechanisms, such as those mediated by IGF2.

genetics

Enrichment of rare protein truncating variants in amyotrophic lateral sclerosis patients

To discover novel genetic risk factors underlying amyotrophic lateral sclerosis (ALS), we aggregated exomes from 3,864 cases and 7,839 ancestry matched controls. We observed a significant excess of ultra-rare and rare protein-truncating variants (PTV) among ALS cases, which was primarily concentrated in constrained genes; however, a significant enrichment in PTVs does persist in the remaining exome. Through gene level analyses, known ALS genes, SOD1, NEK1, and FUS, were the most strongly associated with disease status. We also observed suggestive statistical evidence for multiple novel genes including DNAJC7, which is a highly constrained gene and a member of the heat shock protein family (HSP40). HSP40 proteins, along with HSP70 proteins, facilitate protein homeostasis, such as folding of newly synthesized polypeptides, and clearance of degraded proteins. When these processes are not regulated, misfolding and accumulation of degraded proteins can occur leading to aberrant protein aggregation, one of the pathological hallmarks of neurodegeneration.

genomics

GWAS identifies novel risk locus for erectile dysfunction and implicates hypothalamic neurobiology and diabetes in etiology

GWAS of erectile dysfunction (ED) in 6,175 cases among 223,805 European men identified one new locus at 6q16.3 (lead variant rs57989773, OR 1.20 per C-allele; p = 5.71x10-14), located between MCHR2 and SIM1. In-silico analysis suggests SIM1 to confer ED risk through hypothalamic dysregulation; Mendelian randomization indicates genetic risk of type 2 diabetes causes ED. Our findings provide novel insights into the biological underpinnings of ED.

genomics

Common variant burden contributes significantly to the familial aggregation of migraine in 1,589 families

It has long been observed that complex traits, including migraine, often aggregate in families, but the underlying genetic architecture behind this is not well understood. Two competing hypotheses exist, emphasizing either rare or common genetic variation. More specifically, familial aggregation could be predominantly explained by rare, penetrant variants that segregate according to Mendelian inheritance or rather by the sufficient polygenic accumulation of many common variants, each with an individually small effect. Some combination of both common and rare variation could also contribute towards a spectrum of disease risk.\n\nWe investigated this in a collection of 8,319 individuals across 1,589 migraine families from Finland. Family members were individually diagnosed by a migraine-specific questionnaire with either migraine without aura (MO, ICHD-3 code 1.1, n=2,357), migraine with typical aura (ICHD- 3 code 1.2.1, n=2,420), hemiplegic migraine (HM, ICHD-3 code 1.2.3, n=540), or no migraine (n=3,002). For comparison, we used population-based migraine cases (n=1,101) and controls (n=13,369) from the FINRISK study. The disease status of FINRISK individuals was assigned based on health registry data from outpatient clinics and/or prescription medication. All individuals were genotyped on the Illumina(R) CoreExome or PsychArray chip platforms and imputed to a Finnish reference panel of 6,962 haplotypes. Polygenic risk scores (PRS), representing the common variant burden in each individual, were calculated using weights from the most recent large-scale genome-wide association study of migraine. To account for family structure in our analyses, we used a mixed-model approach, adjusting for the genetic relationship matrix as a random effect.\n\nWe found a significantly higher common variant burden in familial cases of migraine (for all subtypes, measured by the odds ratio [OR] per standard deviation [SD] increase in PRS; OR = 1.76, 95% CI = 1.71-1.81, P = 1.7x10-109) compared to cases from a population cohort (OR = 1.32, 95% CI = 1.25-1.38, P = 7.2x10-17) when using the population controls as a reference group. The highest enrichment was observed for HM (OR = 1.96, 95% CI = 1.86-2.07, P = 8.7x10-36) and migraine with typical aura (OR = 1.85, 95% CI = 1.79-1.91, P = 1.4x10-86) but enrichment was also present for MO (OR = 1.57, 95% CI = 1.51-1.63, P = 1.1x10-48). Comparing within cases, there was no significant difference in common variant burden between the migraine with aura subtypes, HM and migraine with typical aura (OR = 1.09, 95% CI = 0.99-1.19, P = 0.09), but both showed significantly higher enrichment compared to MO (OR = 1.28, 95% CI = 1.17-1.38, P = 7.3x10-7, and OR = 1.17, 95% CI = 1.11-1.23, P = 4.62x10-5, respectively). Additionally, we found that higher common variant burden corresponded to earlier age of headache onset (OR per SD increase in PRS for 3,631 cases with onset before 20 years old compared to 1,686 cases with onset later than 20 years old; OR = 1.11, 95% CI = 1.05-1.18, P = 8.3x10-4). FINRISK population cases identified from national health registry data were found to have lower common variant burden in comparison to the familial migraine cases (OR = 1.32, 95% CI = 1.25-1.38, P = 6.8x10-17), unless the individuals had attended both a specialist clinic and also received prophylactic migraine treatment (OR = 1.70, 95% CI = 1.53-1.88, P = 3.9x10-9). Finally, although rare variants have been suggested as the primary cause for familial hemiplegic migraine (FHM), we found only four out of 45 sequenced FHM families (8.9%) with a pathogenic mutation in one of the known risk genes.\n\nIn summary, our results demonstrate a substantial contribution of common polygenic variation to familial aggregation in migraine, comparable to both controls and that observed in migraine cases from a population cohort. The findings also suggest that individuals with migraine aura symptoms (either typical aura, which is mostly visual, or rare motor aura) tend to have higher common variant burden on average supporting the polygenic model also in these migraine subtypes.

genetics

New synthetic-diploid benchmark for accurate variant calling evaluation

Constructed from the consensus of multiple variant callers based on short-read data, existing benchmark datasets for evaluating variant calling accuracy are biased toward easy regions accessible by known algorithms. We derived a new benchmark dataset from the de novo PacBio assemblies of two human cell lines that are homozygous across the whole genome. This benchmark provides a more accurate and less biased estimate of the error rate of small variant calls in a realistic context.

bioinformatics

Scaling accurate genetic variant discovery to tens of thousands of samples

Comprehensive disease gene discovery in both common and rare diseases will require the efficient and accurate detection of all classes of genetic variation across tens to hundreds of thousands of human samples. We describe here a novel assembly-based approach to variant calling, the GATK HaplotypeCaller (HC) and Reference Confidence Model (RCM), that determines genotype likelihoods independently per-sample but performs joint calling across all samples within a project simultaneously. We show by calling over 90,000 samples from the Exome Aggregation Consortium (ExAC) that, in contrast to other algorithms, the HC-RCM scales efficiently to very large sample sizes without loss in accuracy; and that the accuracy of indel variant calling is superior in comparison to other algorithms. More importantly, the HC-RCM produces a fully squared-off matrix of genotypes across all samples at every genomic position being investigated. The HC-RCM is a novel, scalable, assembly-based algorithm with abundant applications for population genetics and clinical studies.

genomics

Widespread pleiotropy confounds causal relationships between complex traits and diseases inferred from Mendelian randomization

A fundamental assumption in inferring causality of an exposure on complex disease using Mendelian randomization (MR) is that the genetic variant used as the instrumental variable cannot have pleiotropic effects. Violation of this no pleiotropy assumption can cause severe bias. Emerging evidence have supported a role for pleiotropy amongst disease-associated loci identified from GWA studies. However, the impact and extent of pleiotropy on MR is poorly understood. Here, we introduce a method called the Mendelian Randomization Pleiotropy RESidual Sum and Outlier (MR-PRESSO) test to detect and correct for pleiotropy in multi-instrument summary-level MR testing. We show using simulations that existing approaches are less sensitive to the detection of pleiotropy when it occurs in a subset of instrumental variables, as compared to MR-PRESSO. Next, we show that pleiotropy is widespread in MR, occurring in 41% amongst significant causal relationships (out of 4,250 MR tests total) from pairwise comparisons of 82 complex traits and diseases from summary level genome-wide association data. We demonstrate that pleiotropy causes distortion between-168% and 189% of the causal estimate in MR. Furthermore, pleiotropy induces false positive causal relationships-defined as those causal estimates that were no longer statistically significant in the pleiotropy corrected MR test but were previously significant in the naive MR test-in up to 10% of the MR tests using a P < 0.05 cutoff that is commonly used in MR studies. Finally, we show that MR-PRESSO can correct for distortion in the causal estimate in most cases. Our results demonstrate that pleiotropy is widespread and pervasive, and must be properly corrected for in order to maintain the validity of MR.

genomics

Quantifying the impact of rare and ultra-rare coding variation across the phenotypic spectrum

There is a limited understanding about the impact of rare protein truncating variants across multiple phenotypes. We explore the impact of this class of variants on 13 quantitative traits and 10 diseases using whole-exome sequencing data from 100,296 individuals. Protein truncating variants in genes intolerant to this class of mutations increased risk of autism, schizophrenia, bipolar disorder, intellectual disability, ADHD. In individuals without these disorders, there was an association with shorter height, lower education, increased hospitalization and reduced age. Gene sets implicated from GWAS did not show a significant protein truncating variants-burden beyond what captured by established Mendelian genes. In conclusion, we provide the most thorough investigation to date of the impact of rare deleterious coding variants on complex traits, suggesting widespread pleiotropic risk.\n\nMain abbreviations

genetics

The iPSYCH2012 case-cohort sample: New directions for unravelling genetic and environmental architectures of severe mental disorders

The iPSYCH consortium has established a large Danish population-based Case-Cohort sample (iPSYCH2012) aimed at unravelling the genetic and environmental architecture of severe mental disorders. The iPSYCH2012 sample is nested within the entire Danish population born 1981-2005 including 1,472,762 persons. This paper introduces the iPSYCH2012 sample and outlines key future research directions. Cases were identified as persons with schizophrenia (N=3,540), autism (N=16,146), ADHD (N=18,726), and affective disorder (N=26,380), of which 1928 had bipolar affective disorder. Controls were randomly sampled individuals (N=30,000). Within the sample of 86,189 individuals, a total of 57,377 individuals had at least one major mental disorder. DNA was extracted from the neonatal dried blood spot samples obtained from the Danish Neonatal Screening Biobank and genotyped using the Illumina PsychChip. Genotyping was successful for 90% of the sample. The assessments of exome sequencing, methylation profiling, metabolome profiling, vitamin-D, inflammatory and neurotrophic factors are in progress. For each individual, the iPSYCH2012 sample also includes longitudinal information on health, prescribed medicine, social and socioeconomic information and analogous information among relatives. To the best of our knowledge, the iPSYCH2012 sample is the largest and most comprehensive data source for the combined study of genetic and environmental aetiologies of severe mental disorders.

genetics

MTAG: Multi-Trait Analysis of GWAS

We introduce Multi-Trait Analysis of GWAS (MTAG), a method for joint analysis of summary statistics from GWASs of different traits, possibly from overlapping samples. We apply MTAG to summary statistics for depressive symptoms (Neff = 354,862), neuroticism (N = 168,105), and subjective well-being (N = 388,538). Compared to 32, 9, and 13 genome-wide significant loci in the single-trait GWASs (most of which are themselves novel), MTAG increases the number of loci to 64, 37, and 49, respectively. Moreover, association statistics from MTAG yield more informative bioinformatics analyses and increase variance explained by polygenic scores by approximately 25%, matching theoretical expectations.

genomics

Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types

Genetics can provide a systematic approach to discovering the tissues and cell types relevant for a complex disease or trait. Identifying these tissues and cell types is critical for following up on non-coding allelic function, developing ex-vivo models, and identifying therapeutic targets. Here, we analyze gene expression data from several sources, including the GTEx and PsychENCODE consortia, together with genome-wide association study (GWAS) summary statistics for 48 diseases and traits with an average sample size of 169,331, to identify disease-relevant tissues and cell types. We develop and apply an approach that uses stratified LD score regression to test whether disease heritability is enriched in regions surrounding genes with the highest specific expression in a given tissue. We detect tissue-specific enrichments at FDR < 5% for 34 diseases and traits across a broad range of tissues that recapitulate known biology. In our analysis of traits with observed central nervous system enrichment, we detect an enrichment of neurons over other brain cell types for several brain-related traits, enrichment of inhibitory over excitatory neurons for bipolar disorder but excitatory over inhibitory neurons for schizophrenia and body mass index, and enrichments in the cortex for schizophrenia and in the striatum for migraine. In our analysis of traits with observed immunological enrichment, we identify enrichments of T cells for asthma and eczema, B cells for primary biliary cirrhosis, and myeloid cells for Alzheimer's disease, which we validated with independent chromatin data. Our results demonstrate that our polygenic approach is a powerful way to leverage gene expression data for interpreting GWAS signal.

genetics