bioRxiv ScienceSearch

Biology subjects

Johnson, W. C.

Publications and source records attributed to Johnson, W. C..

3 recordsLinked to original sources

Epigenome-wide association analysis of daytime sleepiness in the Multi-Ethnic Study of Atherosclerosis reveals African-American specific associations

Study ObjectivesExcessive daytime sleepiness (EDS) is a consequence of inadequate sleep, or of a primary disorder of sleep-wake control. Population variability in prevalence of EDS and susceptibility to EDS are likely due to genetic and biological factors as well as social and environmental influences. Epigenetic modifications (such as DNA methylation-DNAm) are potential influences on a range of health outcomes. Here, we explored the association between DNAm and daytime sleepiness quantified by the Epworth Sleepiness Scale (ESS).\n\nMethodsWe performed multi-ethnic and ethnic-specific epigenome-wide association studies for DNAm and ESS in 619 individuals from the Multi-Ethnic Study of Atherosclerosis. Replication was assessed in the Cardiovascular Health Study (CHS). Genetic variants in genes proximal to ESS-associated DNAm were analyzed to identify methylation quantitative trait loci and followed with replication of genotype-sleepiness associations in the UK Biobank.\n\nResults61 methylation sites were associated with ESS (FDR [≤] 0.1) in African Americans only, including an association in KCTD5, a gene strongly implicated in sleep. One association (cg26130090) replicated in CHS African Americans (p-value 0.0004). We identified a sleepiness-associated methylation site in the gene RAI1, a gene associated with sleep and circadian phenotypes. In a follow-up analysis, a genetic variant within RAI1 associated with both DNAm and sleepiness score. The variants association with sleepiness was replicated in the UK Biobank.\n\nConclusionsOur analysis identified methylation sites in multiple genes that may be implicated in EDS. These sleepiness-methylation associations were specific to African Americans. Future work is needed to identify mechanisms driving ancestry-specific methylation effects.\n\nStatement of SignificanceExcessive daytime sleepiness is associated with negative health outcomes such as reduction in quality of life, increased workplace accidents, and cardiovascular mortality. There are race/ethnic disparities in excessive daytime sleepiness, however, the environmental and biological mechanisms for these differences are not yet understood. We performed an association analysis of DNA methylation, measured in monocytes, and daytime sleepiness within a racially diverse study population. We detected numerous DNA methylation markers associated with daytime sleepiness in African Americans, but not in European and Hispanic Americans. Future work is required to elucidate the pathways between DNA methylation, sleepiness, and related behavioral/environmental exposures.

genomics

Genetic architecture of gene expression traits across diverse populations

For many complex traits, gene regulation is likely to play a crucial mechanistic role. How the genetic architectures of complex traits vary between populations and subsequent effects on genetic prediction are not well understood, in part due to the historical paucity of GWAS in populations of non-European ancestry. We used data from the MESA (Multi-Ethnic Study of Atherosclerosis) cohort to characterize the genetic architecture of gene expression within and between diverse populations. Genotype and monocyte gene expression were available in individuals with African American (AFA, n=233), Hispanic (HIS, n=352), and European (CAU, n=578) ancestry. We performed expression quantitative trait loci (eQTL) mapping in each population and show genetic correlation of gene expression depends on shared ancestry proportions. Using elastic net modeling with cross validation to optimize genotypic predictors of gene expression in each population, we show the genetic architecture of gene expression for most predictable genes is sparse. We found the best predicted gene, TACSTD2, was the same across populations with R2 > 0.86 in each population. However, we identified a subset of genes that are well-predicted in one population, but poorly predicted in another. We show these differences in predictive performance are due to allele frequency differences between populations. Using genotype weights trained in MESA to predict gene expression in independent populations showed that a training set with ancestry similar to the test set is better at predicting gene expression in test populations, demonstrating an urgent need for diverse population sampling in genomics. Our predictive models and performance statistics in diverse cohorts are made publicly available for use in transcriptome mapping methods at .\n\nAuthor summaryMost genome-wide association studies (GWAS) have been conducted in populations of European ancestry leading to a disparity in understanding the genetics of complex traits between populations. For many complex traits, gene regulation is critical, given the consistent enrichment of regulatory variants among trait-associated variants. However, it is still unknown how the effects of these key variants differ across populations. We used data from MESA to study the underlying genetic architecture of gene expression by optimizing gene expression prediction within and across diverse populations. The populations with genotype and gene expression data available are from individuals with African American (AFA, n=233), Hispanic (HIS, n=352), and European (CAU, n=578) ancestry. After calculating the prediction performance, we found that there are many genes that were well predicted in one population are poorly predicted in another. We further show that a training set with ancestry similar to the test set resulted in better gene expression predictions, demonstrating the need to incorporate diverse populations in genomic studies. Our gene expression prediction models and performance statistics are publicly available to facilitate future transcriptome mapping studies in diverse populations.

genomics

Deep-coverage whole genome sequences and blood lipids among 16,324 individuals

Deep-coverage whole genome sequencing at the population level is now feasible and offers potential advantages for locus discovery, particularly in the analysis rare mutations in non-coding regions. Here, we performed whole genome sequencing in 16,324 participants from four ancestries at mean depth >29X and analyzed correlations of genotypes with four quantitative traits - plasma levels of total cholesterol, low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol, and triglycerides. We conducted a discovery analysis including common or rare variants in coding as well as non-coding regions and developed a framework to interpret genome sequence for dyslipidemia risk. Common variant association yielded loci previously described with the exception of a few variants not captured earlier by arrays or imputation. In coding sequence, rare variant association yielded known Mendelian dyslipidemia genes and, in non-coding sequence, we detected no rare variant association signals after application of four approaches to aggregate variants in non-coding regions. We developed a new, genome-wide polygenic score for LDL-C and observed that a high polygenic score conferred similar effect size to a monogenic mutation (~30 mg/dl higher LDL-C for each); however, among those with extremely high LDL-C, a high polygenic score was considerably more prevalent than a monogenic mutation (23% versus 2% of participants, respectively).

genomics