bioRxiv ScienceSearch

Biology subjects

Zoltan Kutalik

Publications and source records attributed to Zoltan Kutalik.

4 recordsLinked to original sources

Evidence that low socioeconomic position accentuates genetic susceptibility to obesity

Susceptibility to obesity in todays environment has a strong genetic component. Lower socioeconomic position (SEP) is associated with a higher risk of obesity but it is not known if it accentuates genetic susceptibility to obesity. We aimed to use up to 120,000 individuals from the UK Biobank study to test the hypothesis that measures of socioeconomic position accentuate genetic susceptibility to obesity. We used the Townsend deprivation index (TDI) as the main measure of socioeconomic position, and a 69-variant genetic risk score (GRS) as a measure of genetic susceptibility to obesity. We also tested the hypothesis that interactions between BMI genetics and socioeconomic position would result in evidence of interaction with individual measures of the obesogenic environment and behaviours that correlate strongly with socioeconomic position, even if they have no obesogenic role. These measures included self-reported TV watching, diet and physical activity, and an objective measure of activity derived from accelerometers. We performed several negative control tests, including a simulated environment correlated with BMI but not TDI, and sun protection use. We found evidence of gene-environment interactions with TDI (Pinteraction=3x10-10) such that, within the group of 50% living in the most relatively deprived situations, carrying 10 additional BMI-raising alleles was associated with approximately 3.8 kg extra weight in someone 1.73m tall. In contrast, within the group of 50% living in the least deprivation, carrying 10 additional BMI-raising alleles was associated with approximately 2.9 kg extra weight. We also observed evidence of interaction between sun protection use and BMI genetics, suggesting that residual confounding may result in evidence of non-causal interactions. Our findings provide evidence that relative social deprivation best captures aspects of the obesogenic environment that accentuate the genetic predisposition to obesity in the UK.

Genetics

Quantifying the extent to which index event biases influence large genetic association studies

As genetic association studies increase in size to 100,000s of individuals, subtle biases may influence conclusions. One possible bias is \"index event bias\" (IEB), also called \"collider bias\", caused by the stratification by, or enrichment for, disease status when testing associations between gene variants and a disease-associated trait. We first provided a statistical framework for quantifying IEB then identified real examples of IEB in a range of study and analytical designs. We observed evidence of biased associations for some disease alleles and genetic risk scores, even in population-based studies. For example, a genetic risk score consisting of type 2 diabetes variants was associated with lower BMI in 113,203 type 2 diabetes controls from the population based UK Biobank study (-0.010 SDs BMI per allele, P=5E-4), entirely driven by IEB. Three of 11 individual type 2 diabetes risk alleles, and 10 of 25 hypertension alleles were associated with lower BMI at p<0.05 in UK Biobank when analyzing disease free individuals only, of which six hypertension alleles remained associated at p<0.05 after correction for IEB. Our analysis suggested that the associations between CCND2 and TCF7L2 diabetes risk alleles and BMI could (at least partially) be explained by IEB. Variants remaining associated after correction may be pleiotropic and include those in CYP17A1 (allele associated with hypertension risk and lower BMI). In conclusion, IEB may result in false positive or negative associations in very large studies stratified or strongly enriched for/against disease cases.

Genetics

Genome-wide association between transcription factor expression and chromatin accessibility reveals chromatin state regulators

To better understand genome regulation, it is important to uncover the role of transcription factors in the process of chromatin structure establishment and maintenance. Here we present a data-driven approach to systematically characterize transcription factors that are relevant for this process. Our method uses a linear mixed modeling approach to combine data sets of transcription factor binding motif enrichments in open chromatin and gene expression across the same set of cell lines. Applying this approach to the ENCODE data set we confirm already known and imply numerous novel transcription factors in playing a role in the establishment or maintenance of open chromatin.

Genomics

Across-cohort QC analyses of genome-wide association study summary statistics from complex traits

Genome-wide association studies (GWASs) have been successful in discovering replicable SNP-trait associations for many quantitative traits and common diseases in humans. Typically the effect sizes of SNP alleles are very small and this has led to large genome-wide association meta-analyses (GWAMA) to maximize statistical power. A trend towards ever-larger GWAMA is likely to continue, yet dealing with summary statistics from hundreds of cohorts increases logistical and quality control problems, including unknown sample overlap, and these can lead to both false positive and false negative findings. In this study we propose a new set of metrics and visualization tools for GWAMA, using summary statistics from cohort-level GWASs. We proposed a pair of methods in examining the concordance between demographic information and summary statistics. In method I, we use the population genetics Fst statistic to verify the genetic origin of each cohort and their geographic location, and demonstrate using GWAMA data from the GIANT Consortium that geographic locations of cohorts can be recovered and outlier cohorts can be detected. In method II, we conduct principal component analysis based on reported allele frequencies, and is able to recover the ancestral information for each cohort. In addition, we propose a new statistic that uses the reported allelic effect sizes and their standard errors to identify significant sample overlap or heterogeneity between pairs of cohorts. Finally, to quantify unknown sample overlap across all pairs of cohorts we propose a method that uses randomly generated genetic predictors that does not require the sharing of individual-level genotype data and does not breach individual privacy.

Genetics