bioRxiv ScienceSearch

Biology subjects

Denny, J. C.

Publications and source records attributed to Denny, J. C..

3 recordsLinked to original sources

Pulling the covers in electronic health records for an association study with self-reported sleep behaviors

The electronic health record (EHR) contains rich histories of clinical care, but has not traditionally been mined for information related to sleep habits. Here we performed a retrospective EHR study and derived a cohort of 3,652 individuals with self-reported sleep behaviors, documented from visits to the sleep clinic. These individuals were obese (mean body mass index 33.6 kg/m2) and had a high prevalence of sleep apnea (60.5%), however we found sleep behaviors largely concordant with prior prospective cohort studies. In our cohort, average wake time was one hour later and average sleep duration was 40 minutes longer on weekends than on weekdays (p<1{middle dot}10-12). Sleep duration also varied considerably as a function of age, and tended to be longer in females and in whites. Additionally, through phenome-wide association analyses, we found an association of long weekend sleep with depression, and an unexpectedly large number of associations of long weekday sleep with mental health and neurological disorders (q<0.05). We then sought to replicate previously published genetic associations with morning/evening preference on a subset of our cohort with extant genotyping data (n=555). While those findings did not replicate in our cohort, a polymorphism (rs3754214) in high linkage disequilibrium with a previously published polymorphism near TARS2 was associated with long sleep duration (p<0.01). Collectively, our results highlight the potential of the EHR for uncovering the correlates of human sleep in real-world populations.

bioinformatics

Using Topic Modeling via Non-negative Matrix Factorization to Identify Relationships between Genetic Variants and Disease Phenotypes: A Case Study of Lipoprotein(a) (LPA)

Genome-wide and phenome-wide association studies are commonly used to identify important relationships between genetic variants and phenotypes. Most of these studies have treated diseases as independent variables and suffered from heavy multiple adjustment burdens due to the large number of genetic variants and disease phenotypes. In this study, we propose using topic modeling via non-negative matrix factorization (NMF) for identifying associations between disease phenotypes and genetic variants. Topic modeling is an unsupervised machine learning approach that can be used to learn the semantic patterns from electronic health record data. We chose rs10455872 in LPA as the predictor since it has been shown to be associated with increased risk of hyperlipidemia and cardiovascular diseases (CVD). Using data of 12,759 individuals from the biobank at Vanderbilt University Medical Center, we trained a topic model using NMF from 1,853 distinct phecodes extracted from the cohorts electronic health records and generated six topics. We quantified their associations with rs10455872 in LPA. Topics indicating CVD had positive correlations with rs10455872 (P < 0.001), replicating a previous finding. We also identified a negative correlation between LPA and a topic representing lung cancer (P < 0.001). Our results demonstrate the applicability of topic modeling in exploring the relationship between the genome and clinical diseases.\n\nAuthor summaryIdentifying the clinical associations of genetic variants remains crucial in understanding how the human genome modulates disease risk. Traditional phenome-wide association studies consider each disease phenotype as an independent variable, however, diseases often present as complex clusters of comorbid conditions. In this study, we propose using topic modeling to model electronic health record data as a mixture of topics (e.g., disease clusters or relevant comorbidities) and testing associations between topics and genetic variants. Our results demonstrated the feasibility of using topic modeling to replicate and discover novel associations between the human genome and clinical diseases.

genetics

An atlas of genetic variation for linking pathogen-induced cellular traits to human disease

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with disease. To facilitate moving from associations to disease mechanisms, we leveraged the role of pathogens in shaping human evolution with the Hi-HOST Phenome Project (H2P2): a catalog of cellular GWAS comprised of 79 phenotypes in response to 8 pathogens in 528 lymphoblastoid cell lines. Seventeen loci surpass genome-wide significance (p<5x10-8) for phenotypes ranging from pathogen replication to cytokine production. Combining H2P2 with clinical association data from the eMERGE Network and experimental validation revealed evidence for mechanisms of action and connections with diseases. We identified a SNP near CXCL10 as a cis-cytokine-QTL and a new risk factor for inflammatory bowel disease. A SNP in ZBTB20 demonstrated pleiotropy, partially mediated through NF{kappa}B signaling, and was associated with viral hepatitis. Data are available in an H2P2 web portal to facilitate further interpreting human genome variation through the lens of cell biology.

genetics