bioRxiv ScienceSearch

Biology subjects

Thornton, T. A.

Publications and source records attributed to Thornton, T. A..

3 recordsLinked to original sources

Genome-Wide Control of Population Structure and Relatedness in Genetic Association Studies via Linear Mixed Models with Orthogonally Partitioned Structure

Linear mixed models (LMMs) have become the standard approach for genetic association testing in the presence of sample structure. However, the performance of LMMs has primarily been evaluated in relatively homogeneous populations of European ancestry, despite many of the recent genetic association studies including samples from worldwide populations with diverse ancestries. In this paper, we demonstrate that existing LMM methods can have systematic miscalibration of association test statistics genome-wide in samples with heterogenous ancestry, resulting in both increased type-I error rates and a loss of power. Furthermore, we show that this miscalibration arises due to varying allele frequency differences across the genome among populations. To overcome this problem, we developed LMM-OPS, an LMM approach which orthogonally partitions diverse genetic structure into two components: distant population structure and recent genetic relatedness. In simulation studies with real and simulated genotype data, we demonstrate that LMM-OPS is appropriately calibrated in the presence of ancestry heterogeneity and outperforms existing LMM approaches, including EMMAX, GCTA, and GEMMA. We conduct a GWAS of white blood cell (WBC) count in an admixed sample of 3,551 Hispanic/Latino American women from the Womens Health Initiative SNP Health Association Resource where LMM-OPS detects genome-wide significant associations with corresponding p-values that are one or more orders of magnitude smaller than those from competing LMM methods. We also identify a genome-wide significant association with regulatory variant rs2814778 in the DARC gene on chromosome 1, which generalizes to Hispanic/Latino Americans a previous association with reduced WBC count identified in African Americans.

genetics

GWAS of QRS Duration Identifies New Loci Specific to Hispanic/Latino Populations

BackgroundThe electrocardiographically quantified QRS duration measures ventricular depolarization and conduction. QRS prolongation has been associated with poor heart failure prognosis and cardiovascular mortality, including sudden death. While previous genome-wide association studies (GWAS) have identified 32 QRS SNPs across 26 loci among European, African, and Asian-descent populations, the genetics of QRS among Hispanics/Latinos has not been previously explored.\n\nMethodsWe performed a GWAS of QRS duration among Hispanic/Latino ancestry populations (n=15,124) from four studies using 1000 Genomes imputed genotype data (adjusted for age, sex, global ancestry, clinical and study-specific covariates). Study-specific results were combined using fixed-effects, inverse variance-weighted meta-analysis.\n\nResultsWe identified six loci associated with QRS (P<5x10-8), including two novel loci: MYOCD, a nuclear protein expressed in the heart, and SYT1, an integral membrane protein. The top association in the MYOCD locus, intronic SNP rs16946539, was found in Hispanics/Latinos with a minor allele frequency (MAF) of 0.04, but is monomorphic in European and African descent populations. The most significant QRS duration association was for intronic SNP rs3922344 (P= 8.56x10-26) in SCN5A/SCN10A. Three additional previously identified loci, CDKN1A, VTI1A, and HAND1, also exceeded the GWAS significance threshold among Hispanics/Latinos. A total of 27 of 32 previously identified QRS duration SNPs were shown to generalize in Hispanics/Latinos.\n\nConclusionsOur QRS duration GWAS, the first in Hispanic/Latino populations, identified two new loci, underscoring the utility of extending large scale genomic studies to currently under-examined populations.

genetics

Generalizing Genetic Risk Scores from Europeans to Hispanics/Latinos

Genetic risk scores (GRSs) are weighted sums of risk allele counts of single nucleotide polymorphisms (SNPs) associated with a disease or trait. Construction of GRSs is typically based on published results from Genome-Wide Association Studies (GWASs), the majority of which have been performed in large populations of European ancestry (EA) individuals. While many genotype-trait associations have been shown to generalize from EA populations to other populations, such as Hispanics/Latinos, the optimal choice of SNPs and weights for GRSs may differ between populations due to different linkage disequilibrium (LD) and allele frequency patterns. This is further complicated by the fact that different Hispanic/Latino populations may have different admixture patterns, so that LD and allele frequency patterns may not be the same among non-EA populations. Here, we compare various approaches for GRS construction, using GWAS results from both large EA studies and a smaller study in Hispanics/Latinos, the Hispanic Community Health Study/Study of Latinos (HCHS/SOL, n = 12, 803). We consider multiple ways to select SNPs from association regions and to calculate the SNP weights. We study the performance of the resulting GRSs in an independent study of Hispanics/Latinos from the Woman Health Initiative (WHI, n = 3, 582). We support our investigation with simulation studies of potential genetic architectures in a single locus. We observed that selecting variants based on EA GWASs generally performs well, as long as SNP weights are calculated using Hispanics/Latinos GWASs, or using the meta-analysis of EA and Hispanics/Latinos GWASs. The optimal approach depends on the genetic architecture of the trait.

genetics