bioRxiv ScienceSearch

Biology subjects

Ritchie, M. D.

Publications and source records attributed to Ritchie, M. D..

5 recordsLinked to original sources

Collective feature selection to identify crucial epistatic variants

BackgroundMachine learning methods have gained popularity and practicality in identifying linear and non-linear effects of variants associated with complex disease/traits. Detection of epistatic interactions still remains a challenge due to the large number of features and relatively small sample size as input, thus leading to the so-called \"short fat data\" problem. The efficiency of machine learning methods can be increased by limiting the number of input features. Thus, it is very important to perform variable selection before searching for epistasis. Many methods have been evaluated and proposed to perform feature selection, but no single method works best in all scenarios. We demonstrate this by conducting two separate simulation analyses to evaluate the proposed collective feature selection approach.\n\nResultsThrough our simulation study we propose a collective feature selection approach to select features that are in the \"union\" of the best performing methods. We explored various parametric, non-parametric, and data mining approaches to perform feature selection. We choose our top performing methods to select the union of the resulting variables based on a user-defined percentage of variants selected from each method to take to downstream analysis. Our simulation analysis shows that non-parametric data mining approaches, such as MDR, may work best under one simulation criteria for the high effect size (penetrance) datasets, while non-parametric methods designed for feature selection, such as Ranger and Gradient boosting, work best under other simulation criteria. Thus, using a collective approach proves to be more beneficial for selecting variables with epistatic effects also in low effect size datasets and different genetic architectures. Following this, we applied our proposed collective feature selection approach to select the top 1% of variables to identify potential interacting variables associated with Body Mass Index (BMI) in ~44,000 samples obtained from Geisingers MyCode Community Health Initiative (on behalf of DiscovEHR collaboration).\n\nConclusionsIn this study, we were able to show that selecting variables using a collective feature selection approach could help in selecting true positive epistatic variables more frequently than applying any single method for feature selection via simulation studies. We were able to demonstrate the effectiveness of collective feature selection along with a comparison of many methods in our simulation analysis. We also applied our method to identify non-linear networks associated with obesity.

bioinformatics

Genome-wide association meta-analysis of PR interval identifies 47 novel loci associated with atrial and atrioventricular electrical activity

Electrocardiographic PR interval measures atrial and atrioventricular depolarization and conduction, and abnormal PR interval is a risk factor for atrial fibrillation and heart block. We performed a genome-wide association study in over 92,000 individuals of European descent and identified 44 loci associated with PR interval (34 novel). Examination of the 44 loci revealed known and novel biological processes involved in cardiac atrial electrical activity, and genes in these loci were highly over-represented in several cardiac disease processes. Nearly half of the 61 independent index variants in the 44 loci were associated with atrial or blood transcript expression levels, or were in high linkage disequilibrium with one or more missense variants. Cardiac regulatory regions of the genome as measured by cardiac DNA hypersensitivity sites were enriched for variants associated with PR interval, compared to non-cardiac regulatory regions. Joint analyses combining PR interval with heart rate, QRS interval, and atrial fibrillation identified additional new pleiotropic loci. The majority of associations discovered in European-descent populations were also present in African-American populations. Meta-analysis examining over 105,000 individuals of African and European descent identified additional novel PR loci. These additional analyses identified another 13 novel loci. Together, these findings underscore the power of GWAS to extend knowledge of the molecular underpinnings of clinical processes.

genetics

Profiling copy number variation and disease associations from 50,726 DiscovEHR Study exomes

Copy number variants (CNVs) are a substantial source of genomic variation and contribute to a wide range of human disorders. Gene-disrupting exonic CNVs have important clinical implications as they can underlie variability in disease presentation and susceptibility. The relationship between exonic CNVs and clinical traits has not been broadly explored at the population level, primarily due to technical challenges. We surveyed common and rare CNVs in the exome sequences of 50,726 adult DiscovEHR study participants with linked electronic health records (EHRs). We evaluated the diagnostic yield and clinical expressivity of known pathogenic CNVs, and performed tests of association with EHR-derived serum lipids, thereby evaluating the relationship between CNVs and complex traits and phenotypes in an unbiased, real-world clinical context. We identified CNVs from megabase to exon-level resolution, demonstrating reliable, high-throughput detection of clinically relevant exonic CNVs. In doing so, we created a catalog of high-confidence common and rare CNVs and refined population frequency estimates of known and novel gene-disrupting CNVs. Our survey among an unselected clinical population provides further evidence that neuropathy-associated duplications and deletions in 17p12 have similar population prevalence but are clinically under-diagnosed. Similarly, adults who harbor 22q11.2 deletions frequently had EHR documentation of neurodevelopmental/neuropsychiatric disorders and congenital anomalies, but not a formal genetic diagnosis (i.e., deletion). In an exome-wide association study of lipid levels, we identified a novel five-exon duplication within LDLR segregating in a large kindred with features of familial hypercholesterolemia. Exonic CNVs provide new opportunities to understand and diagnose human disease.

genomics

A simulation study investigating power estimates in Phenome-Wide Association Studies

BackgroundPhenome-wide association studies (PheWAS) are a high-throughput approach to evaluate comprehensive associations between genetic variants and a wide range of phenotypic measures. PheWAS has varying sample sizes for quantitative traits, and variable numbers of cases and controls for binary traits across the many phenotypes of interest, which can affect the statistical power to detect associations. The motivation of this study is to investigate the various parameters which affect the estimation of statistical power in PheWAS, including sample size, case-control ratio, minor allele frequency, and disease penetrance.\n\nResultsWe performed a PheWAS simulation study, where we investigated variations in statistical power based on different parameters, such as overall sample size, number of cases, case-control ratio, minor allele frequency, and disease penetrance. The simulation was performed on both binary and quantitative phenotypic measures. Our simulation on binary traits suggests that the number of cases has more impact than the case to control ratio; also, we found that a sample size of 200 cases or more maintains the statistical power to identify associations for common variants. For quantitative traits, a sample size of 1000 or more individuals performed best in the power calculations. We focused on common genetic variants (MAF>0.01) in this study; however, in future studies, we will be extending this effort to perform similar simulations on rare variants.\n\nConclusionsThis study provides a series of PheWAS simulation analyses that can be used to estimate statistical power for some potential scenarios. These results can be used to provide guidelines for appropriate study design for future PheWAS analyses.

genomics

Depression Linked to Frequent Emergency Department Use in Large 10-year Retrospective Analysis of an Integrated Health Care System

We evaluated general patient features related to depression and frequency of Emergency Department (ED) use in a large integrated health care system. Electronic Health Records of 287,281 adults from a general patient population were studied retrospectively over a 10-year period. Patients with a history of depression were more likely to be seen in the ED and at higher frequency than those without. Frequent ED users were more likely to have a history of depression or psychiatric medication orders than infrequent users. ED visits by depression patients and frequent users have highly correlated complaints and discharge diagnoses with other ED users, often related to pain. Poorly managed depression may be playing a role in frequent ED utilization which may be addressed by universal screening for depression, evaluation of barriers to treatment, and other novel interventions to improve care coordination.

epidemiology