bioRxiv ScienceSearch

Biology subjects

Laurie, C. C.

Publications and source records attributed to Laurie, C. C..

4 recordsLinked to original sources

GWAS of QRS Duration Identifies New Loci Specific to Hispanic/Latino Populations

BackgroundThe electrocardiographically quantified QRS duration measures ventricular depolarization and conduction. QRS prolongation has been associated with poor heart failure prognosis and cardiovascular mortality, including sudden death. While previous genome-wide association studies (GWAS) have identified 32 QRS SNPs across 26 loci among European, African, and Asian-descent populations, the genetics of QRS among Hispanics/Latinos has not been previously explored.\n\nMethodsWe performed a GWAS of QRS duration among Hispanic/Latino ancestry populations (n=15,124) from four studies using 1000 Genomes imputed genotype data (adjusted for age, sex, global ancestry, clinical and study-specific covariates). Study-specific results were combined using fixed-effects, inverse variance-weighted meta-analysis.\n\nResultsWe identified six loci associated with QRS (P<5x10-8), including two novel loci: MYOCD, a nuclear protein expressed in the heart, and SYT1, an integral membrane protein. The top association in the MYOCD locus, intronic SNP rs16946539, was found in Hispanics/Latinos with a minor allele frequency (MAF) of 0.04, but is monomorphic in European and African descent populations. The most significant QRS duration association was for intronic SNP rs3922344 (P= 8.56x10-26) in SCN5A/SCN10A. Three additional previously identified loci, CDKN1A, VTI1A, and HAND1, also exceeded the GWAS significance threshold among Hispanics/Latinos. A total of 27 of 32 previously identified QRS duration SNPs were shown to generalize in Hispanics/Latinos.\n\nConclusionsOur QRS duration GWAS, the first in Hispanic/Latino populations, identified two new loci, underscoring the utility of extending large scale genomic studies to currently under-examined populations.

genetics

A Fully-Adjusted Two-Stage Procedure for Rank Normalization in Genetic Association Studies

When testing genotype-phenotype associations using linear regression, departure of the trait distribution from normality can impact both Type I error rate control and statistical power, with worse consequences for rarer variants. While it has been shown that applying a rank-normalization transformation to trait values before testing may improve these statistical properties, the factor driving them is not the trait distribution itself, but its residual distribution after regression on both covariates and genotype. Because genotype is expected to have a small effect (if any) investigators now routinely use a two-stage method, in which they first regress the trait on covariates, obtain residuals, rank-normalize them, and then secondly use the rank-normalized residuals in association analysis with the genotypes. Potential confounding signals are assumed to be removed at the first stage, so in practice no further adjustment is done in the second stage. Here, we show that this widely-used approach can lead to tests with undesirable statistical properties, due to both a combination of a mis-specified mean-variance relationship, and remaining covariate associations between the rank-normalized residuals and genotypes. We demonstrate these properties theoretically, and also in applications to genome-wide and whole-genome sequencing association studies. We further propose and evaluate an alternative fully-adjusted two-stage approach that adjusts for covariates both when residuals are obtained, and in the subsequent association test. This method can reduce excess Type I errors and improve statistical power.

genetics

Imputation-Based Genomic Coverage Assessments of Current Genotyping Arrays: Illumina HumanCore, OmniExpress, Multi-Ethnic global array and sub-arrays, Global Screening Array, Omni2.5M, Omni5M, and Affymetrix UK Biobank

Genotyping arrays have been widely adopted as an efficient means to interrogate variation across the human genome. Genetic variants may be observed either directly, via genotyping, or indirectly, through linkage disequilibrium with a genotyped variant. The total proportion of genomic variation captured by an array, either directly or indirectly, is referred to as \"genomic coverage.\" Here we use genotype imputation and Phase 3 of the 1000 Genomes Project to assess genomic coverage of several modern genotyping arrays. We find that in general, coverage increases with increasing array density. However, arrays designed to cover specific populations may yield better coverage in those populations compared to denser arrays not tailored to the given population. Ultimately, array choice involves trade-offs between cost, density, and coverage, and our work helps inform investigators weighing these choices and trade-offs.

genetics

Integrated Computing And Tracking System For Centralized High-Throughput Genetic Analysis: A Case Study

The Genetic Analysis Center (GAC) of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) developed an Integrated Computing and Tracking system (ICT) in order to perform genome-wide and other genetic association studies automatically and efficiently, while documenting all analysis specifications. This system provides easy-to-use analysis set-up and computing procedures, automatic reports, and analysis search functionality due to integration with an on-site database. In this paper we describe the ICT and demonstrate how it satisfies key principles of reproducible research, while respecting constraints and challenges arising from using very large, restricted access, human-subjects data. This case study may benefit other groups that have similar requirements for high-throughput analysis execution and management.

bioinformatics