bioRxiv ScienceSearch

Biology subjects

Rice, K. M.

Publications and source records attributed to Rice, K. M..

3 recordsLinked to original sources

Population Stratification at the Phenotypic Variance level and Implication for the Analysis of Whole Genome Sequencing Data from Multiple Studies

In modern Whole Genome Sequencing (WGS) epidemiological studies, participant-level data from multiple studies are often pooled and results are obtained from a single analysis. We consider the impact of differential phenotype variances by study, which we term variance stratification. Unaccounted for, variance stratification can lead to both decreased statistical power, and increased false positives rates, depending on how allele frequencies, sample sizes, and phenotypic variances vary across the studies that are pooled. We describe a WGS-appropriate analysis approach, implemented in freely-available software, which allows study-specific variances and thereby improves performance in practice. We also illustrate the variance stratification problem, its solutions, and a corresponding diagnostic procedure in data from the Trans-Omics for Precision Medicine Whole Genome Sequencing Program (TOPMed), used in association tests for hemoglobin concentrations and BMI.

genetics

Genetic loci associated with prevalent and incident myocardial infarction and coronary heart disease in the Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium

BackgroundGenome-wide association studies have identified multiple genomic loci associated with coronary artery disease, but most are common variants in non-coding regions that provide limited information on causal genes and etiology of the disease. To better understand etiological pathways that might lead to discovery of new treatments or prevention strategies, we focused our investigation on low-frequency and rare sequence variations primarily residing in coding regions of the genome while also exploring associations with common variants. Methods and ResultsUsing samples of individuals of European ancestry from ten cohorts within the Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) consortium, both cross-sectional and prospective analyses were conducted to examine associations between genetic variants and myocardial infarction (MI), coronary heart disease (CHD), and all-cause mortality following these events. Single variant and gene-based analyses were performed separately in each cohort and then meta-analyzed for each outcome. A low-frequency intronic variant (rs988583) in PLCL1 was significantly associated with prevalent MI (OR=1.80, 95% confidence interval: 1.43, 2.27; P=7.12 x 10-7). Three common variants, rs9349379 in PHACTR1, and rs1333048 and rs4977574 in the 9p21 region, were significantly associated with prevalent CHD. Four common variants (rs4977574, rs10757278, rs1333049, and rs1333048) within the 9p21 locus were significantly associated with incident MI. We conducted gene-based burden tests for genes with a cumulative minor allele count (cMAC) [&ge;] 5 and variants with minor allele frequency (MAF) < 5%. TMPRSS5 and LDLRAD1 were significantly associated with prevalent MI and CHD, respectively, and RC3H2 and ANGPTL4 were significantly associated with incident MI and CHD, respectively. No loci were significantly associated with all-cause mortality following a MI or CHD event. ConclusionThis study confirmed previously reported loci influencing heart disease risk, and one single variant and three genes associated with MI and CHD were newly identified and warrant future investigation.

genetics

A Fully-Adjusted Two-Stage Procedure for Rank Normalization in Genetic Association Studies

When testing genotype-phenotype associations using linear regression, departure of the trait distribution from normality can impact both Type I error rate control and statistical power, with worse consequences for rarer variants. While it has been shown that applying a rank-normalization transformation to trait values before testing may improve these statistical properties, the factor driving them is not the trait distribution itself, but its residual distribution after regression on both covariates and genotype. Because genotype is expected to have a small effect (if any) investigators now routinely use a two-stage method, in which they first regress the trait on covariates, obtain residuals, rank-normalize them, and then secondly use the rank-normalized residuals in association analysis with the genotypes. Potential confounding signals are assumed to be removed at the first stage, so in practice no further adjustment is done in the second stage. Here, we show that this widely-used approach can lead to tests with undesirable statistical properties, due to both a combination of a mis-specified mean-variance relationship, and remaining covariate associations between the rank-normalized residuals and genotypes. We demonstrate these properties theoretically, and also in applications to genome-wide and whole-genome sequencing association studies. We further propose and evaluate an alternative fully-adjusted two-stage approach that adjusts for covariates both when residuals are obtained, and in the subsequent association test. This method can reduce excess Type I errors and improve statistical power.

genetics