bioRxiv ScienceSearch

Biology subjects

Ober, C.

Publications and source records attributed to Ober, C..

9 recordsLinked to original sources

Shared and distinct genetic risk factors for childhood onset and adult onset asthma

BackgroundChildhood and adult onset asthma differ with respect to severity and co-morbidities. Whether they also differ with respect to genetic risk factors has not been previously investigated.\n\nMethodsWe used data from the UK Biobank to conduct genome-wide association studies (GWASs) in 9,433 childhood onset asthma (onset before age 12) and 21,564 adult onset asthma (onset between ages 26 and 65) cases, each compared to 318,237 non-asthmatic controls (older than age 38), and for age of onset in 37,846 asthma cases. Enrichment studies determined the tissues in which genes at GWAS loci were most highly expressed, and PrediXcan, a transcriptome-wide gene-based test, was used to identify candidate risk genes.\n\nFindingsWe detected 61 independent asthma loci: 23 were childhood onset specific, one was adult onset specific, and 37 were shared. Nineteen loci were associated with age of asthma onset. Genes at the childhood onset loci were most highly expressed in skin, blood and small intestine; genes at the adult onset loci were most highly expressed in lung, blood, small intestine and spleen. PrediXcan identified 113 unique candidate genes at 22 of the 61 GWAS loci.\n\nInterpretationGenetic risk factors for adult onset asthma are largely a subset of the genetic risk for childhood onset asthma but with overall smaller effects, suggesting a greater role for non-genetic risk factors in adult onset asthma. In contrast, the onset of disease in childhood is associated with additional genes with relatively large effect sizes. Combined with gene expression and tissue enrichment patterns, we suggest that the establishment of disease in children is driven more by allergy and epithelial barrier dysfunction whereas the etiology of adult onset asthma is more lung-centered, with immune mediated pathways driving disease progression in both children and adults.\n\nFundingThis work was supported by the National Institutes of Health grants R01 MH107666 and P30 DK20595 to H.K.I., R01 HL129735, R01 HL122712, P01 HL070831, and UG3 OD023282 to C.O.; N.S. was supported by T32 HL007605.

genetics

Heritability Estimation and Differential Analysis with Generalized Linear Mixed Models in Genomic Sequencing Studies

MotivationGenomic sequencing studies, including RNA sequencing and bisulfite sequencing studies, are becoming increasingly common and increasingly large. Large genomic sequencing studies open doors for accurate molecular trait heritability estimation and powerful differential analysis. Heritability estimation and differential analysis in sequencing studies requires the development of statistical methods that can properly account for the count nature of the sequencing data and that are computationally efficient for large data sets.\n\nResultsHere, we develop such a method, PQLseq (Penalized Quasi-Likelihood for sequencing count data), to enable effective and efficient heritability estimation and differential analysis using the generalized linear mixed model framework. With extensive simulations and comparisons to previous methods, we show that PQLseq is the only method currently available that can produce unbiased heritability estimates for sequencing count data. In addition, we show that PQLseq is well suited for differential analysis in large sequencing studies, providing calibrated type I error control and more power compared to the standard linear mixed model methods. Finally, we apply PQLseq to perform gene expression heritability estimation and differential expression analysis in a large RNA sequencing study in the Hutterites.\n\nAvailability and implementationPQLseq is implemented as an R package with source code freely available at www.xzlab.org/software.html and https://cran.r-project.org/web/packages/PQLseq/index.html.\n\nContactXZ (xzhousph@umich.edu)\n\nSupplementary informationSupplementary data are available online.

genomics

Parent of origin gene expression in a founder population identifies two new imprinted genes at known imprinted regions

Genomic imprinting is the phenomena that leads to silencing of one copy of a gene inherited from a specific parent. Mutations in imprinted regions have been involved in diseases showing parent of origin effects. Identifying genes with evidence of parent of origin expression patterns in family studies allows the detection of more subtle imprinting. Here, we use allele specific expression in lymphoblastoid cell lines from 306 Hutterites related in a single pedigree to provide formal evidence for parent of origin effects. We take advantage of phased genotype data to assign parent of origin to RNA-seq reads in individuals with gene expression data. Our approach identified known imprinted genes, two putative novel imprinted genes, and 14 genes with asymmetrical parent of origin gene expression. We used gene expression in peripheral blood leukocytes (PBL) to validate our findings, and then confirmed imprinting control regions (ICRs) using DNA methylation levels in the PBLs.\n\nAuthor SummaryLarge scale gene expression studies have identified known and novel imprinted genes through allele specific expression without knowing the parental origins of each allele. Here, we take advantage of phased genotype data to assign parent of origin to RNA-seq reads in 306 individuals with gene expression data. We identified known imprinted genes as well as two novel imprinted genes in lymphoblastoid cell line gene expression. We used gene expression in PBLs to validate our findings, and DNA methylation levels in PBLs to confirm previously characterized imprinting control regions that could regulate these imprinted genes.

genetics

Longitudinal studies at birth and age 7 reveal strong effects of genetic variation on ancestry-associated DNA methylation patterns in blood cells from ethnically admixed children

Epigenetic architecture is influenced by genetic and environmental factors, but little is known about their relative contributions or longitudinal dynamics. Here, we studied DNA methylation (DNAm) at over 750,000 CpG sites in mononuclear blood cells collected at birth and age 7 from 196 children of primarily self-reported Black and Hispanic ethnicities to study race-associated DNAm patterns. We developed a novel Bayesian method for high dimensional longitudinal data and showed that race-associated DNAm patterns at birth and age 7 are nearly identical. Additionally, we estimated that up to 51% of all self-reported race-associated CpGs had race-dependent DNAm levels that were mediated through local genotype and, quite surprisingly, found that genetic factors explained an overwhelming majority of the variation in DNAm levels at other, previously identified, environmentally-associated CpGs. These results not only indicate that race-associated DNAm patterns in blood are present at birth and are primarily genetically, and not environmentally, determined, but also that DNAm in blood cells overall is robust to many environmental exposures during the first 7 years of life.

genomics

Determining the genetic basis of anthracycline-cardiotoxicity by molecular response QTL mapping in induced cardiomyocytes

Anthracycline-induced cardiotoxicity (ACT) is a key limiting factor in setting optimal chemotherapy regimes for cancer patients, with almost half of patients expected to ultimately develop congestive heart failure given high drug doses. However, the genetic basis of sensitivity to anthracyclines such as doxorubicin remains unclear. To begin addressing this, we created a panel of iPSC-derived cardiomyocytes from 45 individuals and performed RNA-seq after 24h exposure to varying levels of doxorubicin. The transcriptomic response to doxorubicin is substantial, with the majority of genes being differentially expressed across treatments of different concentrations and over 6000 genes showing evidence of differential splicing. Overall, our observations indicate that splicing fidelity decreases in the presence of doxorubicin. We detect 376 response-expression QTLs and 42 response-splicing QTLs, i.e. genetic variants that modulate the individual transcriptomic response to doxorubicin in terms of expression and splicing changes respectively. We show that inter-individual variation in transcriptional response is predictive of cell damage measured in vitro using a cardiac troponin assay, which in turn is shown to be associated with in vivo ACT risk. Finally, the molecular QTLs we detected are enriched in lower ACT GWAS p-values, further supporting the in vivo relevance of our map of genetic regulation of cellular response to anthracyclines.

genomics

Gene co-expression networks in whole blood implicate multiple interrelated molecular pathways in obese asthma

BackgroundAsthmatic children who develop obesity have poorer outcomes compared to those that do not, including poorer control, more severe symptoms, and greater resistance to standard treatment. Gene expression networks are powerful statistical tools for characterizing the underpinnings of human disease that leverage the putative co-regulatory relationships of genes to infer biological pathways altered in disease states.\n\nObjectiveThe aim of this study was to characterize the biology of childhood asthma complicated by adult obesity.\n\nMethodsWe performed weighted gene co-expression network analysis (WGCNA) of gene expression data in whole blood from 514 adult subjects from the Childhood Asthma Management Program (CAMP). We then performed module preservation and association replication analyses in 418 subjects from two independent asthma cohorts (one pediatric and one adult).\n\nResultsWe identified a multivariate model in which four gene co-expression network modules were associated with incident obesity in CAMP (each P < 0.05). The module memberships were enriched for genes in pathways related to platelets, integrins, extracellular matrix, smooth muscle, NF-{kappa}B signaling, and Hedgehog signaling. The network structures of each of the four obese asthma modules were significantly preserved in both replication cohorts (permutation P = 9.999E-05). The corresponding module gene sets were significantly enriched for differential expression in obese subjects in both replication cohorts (each P < 0.05).\n\nConclusionsOur gene co-expression network profiles thus implicate multiple interrelated pathways in the biology of an important endotype of obese asthma.\n\nKey MessagesO_LIWe hypothesized that individuals with asthma complicated by obesity had distinct blood gene expression signatures.\nC_LIO_LIGene co-expression network analysis implicated several inflammatory biological pathways in one form of obese asthma.\nC_LI\n\nCapsule SummaryThis work addresses a knowledge gap about the molecular relationship between asthma and obesity, suggesting that an endotype of obese asthma, known as asthma complicated by obesity, is underpinned by coherent biological mechanisms.\n\nAbbreviations

genomics

Parent of Origin Effects on Quantitative Phenotypes in a Founder Population

The impact of the parental origin of associated alleles in GWAS has been largely ignored. Yet sequence variants could affect traits differently depending on whether they are inherited from the mother or the father. To explore this possibility, we studied 21 quantitative phenotypes in a large Hutterite pedigree. We first identified variants with significant single parent (maternal-only or paternal-only) effects, and then used a novel statistical model to identify variants with opposite parental effects. Overall, we identified parent of origin effects (POEs) on 11 phenotypes, most of which are risk factors for cardiovascular disease. Many of the loci with POEs have features of imprinted regions and many of the variants with POE are associated with the expression of nearby genes. Overall, our results indicate that POEs, which are often opposite in direction, are relatively common in humans, have potentially important clinical effects, and will be missed in traditional GWAS.

genetics

Rare non-coding variants are associated with plasma lipid traits in a founder population

Founder populations are ideally suited for studies on the clinical effects of alleles that are rare in general populations but occur at higher frequencies in these isolated populations. Whole genome sequencing in 98 South Dakota Hutterites, a founder population of European descent, and subsequent imputation to the Hutterite pedigree revealed 660,238 single nucleotide polymorphisms (SNPs; 98.9% non-coding) that are rare (<1%) or absent in European populations, but occur at frequencies greater than 1% in the Hutterites. We examined the effects of these rare in European variants on plasma levels of LDL cholesterol (LDL-C), HDL cholesterol (HDL-C), total cholesterol and triglycerides (TG) in 828 Hutterites and applied a Bayesian hierarchical framework to prioritize potentially causal variants based on functional annotations. We identified two novel non-coding rare variants associated with LDL-C (rs17242388 in LDLR) and HDL-C (rs189679427 between GOT2 and APOOP5), and replicated previous associations of a splice variant in APOC3 (rs138326449) with TG and HDL-C. All three variants are at well-replicated loci in genome wide association study (GWAS) but are independent from and have larger effect sizes than the known common variation in these regions. We also identified variants at two novel loci (rs191020975 in EPHA6 and chr1:224811120 in CNIH3) at suggestive levels of significance with LDL-C. Candidate expression quantitative loci (eQTL) analyses in lymphoblastoid cell lines (LCLs) in the Hutterites suggest that these rare non-coding variants are likely to mediate their effects on lipid traits by regulating gene expression. Overall, we provide insights into the mechanisms regulating lipid traits and potentially new therapeutic targets.

genetics

Reducing mitochondrial reads in ATAC-seq using CRISPR/Cas9

ATAC-seq is a high-throughput sequencing technique that aims at identifying DNA sequences located in open chromatin. Depending on the cell type, ATAC-seq may yield a high number of mitochondrial sequencing reads (~20-80% of the reads). As the regions of open chromatin of interest are usually located in the nuclear genome, mitochondrial reads are typically discarded from the analysis. To decrease wasted sequencing, we performed targeted cleavage of mitochondrial DNA using CRISPR/Cas9 and 100 mtDNA-specific guide RNAs. We also tested a modified ATAC-seq protocol that does not include detergent in the cell lysis buffer. Both treatments resulted in considerable reduction of mitochondrial reads (1.7 and 3-fold, respectively). The removal of detergent, however, resulted in increased background and fewer peaks identified. The highest number of peaks and highest quality data was obtained by preparing samples with the original ATAC-seq protocol (using detergent) and treating them with anti-mitochondrial guide RNAs and Cas9. This strategy could lead to considerable cost reduction and improved peak calling when performing ATAC-seq on a moderate to large number of samples and in cell types that contain a large amount of mitochondria.

genomics