bioRxiv Science⌕ Search

Biology subjects

Kolde, A.

Publications and source records attributed to Kolde, A..

3 recordsLinked to original sources

Insights into the metabolic consequences of type 2 diabetes

Circulating metabolite levels have been associated with type 2 diabetes (T2D), but the extent to which these are affected by T2D and the involvement of genetics in mediating these relationships remain to be elucidated. In this study, we investigate the interplay between genetics, metabolomics and T2D risk in the UK Biobank dataset. We find 79 metabolites with a causal association to T2D, mostly spanning lipid-related classes, while twice as many metabolites are causally affected by T2D liability, including branched-chain amino acids. Secondly, using an interaction quantitative trait locus (QTL) analysis, we describe four metabolites, consistently replicated in an independent dataset from the Estonian Biobank, for which genetic loci in two different genomic regions show attenuated regulation in T2D cases compared to controls. The significant variants from the interaction QTL analysis are significant QTLs for the corresponding metabolites in the general population, but are not associated with T2D risk, pointing towards consequences of T2D on the genetic regulation of metabolite levels. Finally, we find 165 metabolites associated with microvascular, macrovascular, or both types of T2D complications, with only a few discriminating between complication classes. Of the 165 metabolites, 40 are not causally linked to T2D in either direction, suggesting biological mechanisms specific to the occurrence of complications. Overall, this work provides a map of the metabolic consequences of T2D and of the genetic regulation of metabolite levels and enable to better understand the trajectory of T2D leading to complications.

genomics↗

Extensive co-regulation of neighbouring genes complicates the use of eQTLs in target gene prioritisation

Identifying causal genes underlying genome-wide association studies (GWAS) is a fundamental problem in human genetics. Although colocalisation with gene expression quantitative trait loci (eQTLs) is often used to prioritise GWAS target genes, systematic benchmarking has been limited due to unavailability of large ground truth datasets. Here, we re-analysed plasma protein QTL data from 3,301 individuals of the INTERVAL cohort together with 131 eQTL Catalogue datasets. Focusing on variants located within or close to the affected protein identified 793 proteins with at least one cis-pQTL where we could assume that the most likely causal gene was the gene coding for the protein. We then benchmarked the ability of cis-eQTLs to recover these causal genes by comparing three Bayesian colocalisation methods (coloc.susie, coloc.abf and CLPP) and five Mendelian randomisation (MR) approaches (three varieties of inverse-variance weighted MR, MR-RAPS, and MRLocus). We found that assigning fine-mapped pQTLs to their closest protein coding genes outperformed all colocalisation methods regarding both precision (71.9%) and recall (76.9%). Furthermore, the colocalisation method with the highest recall (coloc.susie - 46.3%) also had the lowest precision (45.1%). Combining evidence from multiple conditionally distinct colocalising QTLs with MR increased precision to 81%, but this was accompanied by a large reduction in recall to 7.1%. Furthermore, the choice of the MR method greatly affected performance, with the standard inverse-variance weighted MR often producing many false positives. Our results highlight that linking GWAS variants to target genes remains challenging with eQTL evidence alone, and prioritising novel targets requires triangulation of evidence from multiple sources.

genetics↗

Analysis of follow-up data in large biobank cohorts: a review of methodology

This study focuses on key methodological challenges in genome-wide association studies (GWAS) of biobank data with time-to-event outcomes, analyzed using the Cox proportional hazards (CPH) model. We address four primary issues: left-truncation of the data, com-putational inefficiency of standard model-fitting algorithms, related-ness among individuals, and model misspecification. To manage left-truncation, the common practice is to use age as the timescale, with individuals entering the risk set at their age of recruitment. We assess how this choice of timescale influences bias and statistical power, under realistic GWAS conditions of varying effect sizes and censoring rates. In addition, to alleviate the computational burden typical in large-scale data, we propose and evaluate a two-step martingale residual (MR) approach for high-dimensional CPH modeling. Our results show that the timescale choice has minimal effect on accuracy for small hazard ratios, though using birth age as the timescale-ignoring recruitment age-yields the highest power for association detection. We find that relatedness, when ignored, does not substantially bias effect size estimates, while omitting key covariates introduces significant bias. The two-step MR approach proves to be computationally efficient, retaining power for detecting small effect sizes, making it suitable for large-scale association studies. However, when precise effect size estimates are critical, particularly for moderate or larger effect sizes, we recommend recalculating these estimates using the conventional CPH model, with careful attention to left-truncation and relatedness. These conclusions are drawn from simulations and illustrated with data from the Estonian Biobank cohort.

genomics↗