bioRxiv ScienceSearch

Biology subjects

Dey, R.

Publications and source records attributed to Dey, R..

3 recordsLinked to original sources

Genome-wide association study of 1 million people identifies 111 loci for atrial fibrillation

To understand the genetic variation underlying atrial fibrillation (AF), the most common cardiac arrhythmia, we performed a genome-wide association study (GWAS) of > 1 million people, including 60,620 AF cases and 970,216 controls. We identified 163 independent risk variants at 111 loci and prioritized 165 candidate genes likely to be involved in AF. Many of the identified risk variants fall near genes where more deleterious mutations have been reported to cause serious heart defects in humans or mice (MYH6, NKX2-5, PITX2, TBC1D32, TBX5),1,2 or near genes important for striated muscle function and integrity (e.g. MYH7, PKP2, SSPN, SGCA). Experiments in rabbits with heart failure and left atrial dilation identified a heterogeneous distributed molecular switch from MYH6 to MYH7 in the left atrium, which resulted in contractile and functional heterogeneity and may predispose to initiation and maintenance of atrial arrhythmia.

genetics

Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies

In genome-wide association studies (GWAS) for thousands of phenotypes in large biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, linear mixed model and the recently proposed logistic mixed model, perform poorly - producing large type I error rates - in the analysis of phenotypes with unbalanced case-control ratios. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approximation (SPA) to calibrate the distribution of score test statistics. This method, SAIGE, provides accurate p-values even when case-control ratios are extremely unbalanced. It utilizes state-of-art optimization strategies to reduce computational time and memory cost of generalized mixed model. The computation cost linearly depends on sample size, and hence can be applicable to GWAS for thousands of phenotypes by large biobanks. Through the analysis of UK Biobank data of 408,961 white British European-ancestry samples for >1400 binary phenotypes, we show that SAIGE can efficiently analyze large sample data, controlling for unbalanced case-control ratios and sample relatedness.

genomics

A fast and accurate algorithm to test for binary phenotypes and its application to PheWAS

The availability of electronic health record (EHR)-based phenotypes allows for genome-wide association analyses in thousands of traits, and has great potential to identify novel genetic variants associated with clinical phenotypes. We can interpret the phenome-wide association study (PheWAS) result for a single genetic variant by observing its association across a landscape of phenotypes. Since PheWAS can test 1000s of binary phenotypes, and most of them have unbalanced (case:control = 1:10) or often extremely unbalanced (case:control = 1:600) case-control ratios, existing methods cannot provide an accurate and scalable way to test for associations. Here we propose a computationally fast score test-based method that estimates the distribution of the test statistic using the saddlepoint approximation. Our method is much faster than the state of the art Firths test ([~] 100 times). It can also adjust for covariates and control type I error rates even when the case-control ratio is extremely unbalanced. Through application to PheWAS data from the Michigan Genomics Initiative, we show that the proposed method can control type I error rates while replicating previously known association signals even for traits with a very small number of cases and a large number of controls.

genomics