bioRxiv Science⌕ Search

Biology subjects

Masala, M.

Publications and source records attributed to Masala, M..

4 recordsLinked to original sources

COLSTATS: an interactive web platform for systematic colocalization of GWAS and QTL summary statistics

Colocalization of GWAS with QTL data is pivotal for identifying shared genetic determinants of molecular traits and diseases yet remains hindered by fragmented data sources and technical barriers. COLSTATS addresses this gap by offering a web-based Shiny platform that integrates 1,648,851 harmonized summary statistic datasets from key resources (GWAS Catalog, UK Biobank, GTEx, BluePrint, eQTLGen, and immune-phenotype studies), all formatted as VCFs aligned to hg38 with uniform chr_pos_ref_alt variant identifiers. Users can interactively select two traits, specify a genomic region or gene, define priors, and execute colocalization via coloc.abf, with results displayed in tabular and graphical form, downloadable as PDF, CSV, or PNG. The system includes result caching, history tracking, and an intuitive interface for streamlined exploration.COLSTATS is freely accessible at https://colstats.irgb.cnr.it. This resource empowers researchers, including non-bioinformaticians, to perform transparent and reproducible colocalization analyses efficiently.

bioinformatics↗

Quantifying Uncertainty in Polygenic Risk Scores Using Conformalized Quantile Regression

Polygenic risk scores (PRS) are widely used in post-GWAS analyses to predict complex traits across humans, animals, and plants, yet the uncertainty of these predictions is rarely quantified at the individual level. Here, we introduce a framework for individualized uncertainty quantification based on quantile regression and conformal prediction, enabling the construction of prediction intervals with guaranteed coverage under minimal assumptions. Quantile regression enables adaptive, individual-specific prediction intervals that capture asymmetry and allow interval widths to vary substantially across individuals based on genetic information alone. Applying this framework to 62 traits in the UK Biobank and the ProgeNIA/SardiNIA studies, we show that these intervals maintain valid coverage and reduce uncertainty in risk stratification compared to existing methods, driven by their adaptive construction. Prediction interval width correlates positively with age and BMI, indicating reduced genetic predictability in subsets of the population where genetic effects interact with environmental factors. Our results demonstrate that incorporating uncertainty is essential for interpreting polygenic predictions and provide a principled approach to distinguish individuals whose phenotypes are well explained by genetic predictors from those in whom non-genetic influences dominate.

bioinformatics↗

Regenie.QRS: computationally efficient whole-genome quantile regression at biobank scale

Genotype-phenotype associations can be context-dependent and dynamic in nature leading to heterogeneity of genetic effects across different parts of the phenotype distribution. Quantile regression, an alternative to linear regression for continuous phenotypes, is particularly well suited for detecting and characterizing heterogeneous genotype-phenotype associations. Here we propose a novel and computationally efficient whole-genome quantile regression technique, Regenie.QRS, for biobank-scale GWAS data with genetic structure. Our approach first estimates the polygenic effect, and then incorporates this effect as an offset in the non-mixed quantile regression model. Our simulations demonstrate robust control of type I error and higher power to detect heterogeneous associations relative to linear regression in GWAS, and improved power over the marginal quantile regression tests. We present applications using data from the UK Biobank and the ProgeNIA/SardiNIA project, where we show the advantages of Regenie.QRS in identifying and characterizing heterogeneous genetic effects. To cite just one interesting example, using quantile regression we are able to show that even though variants at the G6PC2 locus increase glucose levels, their effects are much stronger at lower quantiles of glucose level distribution than at higher quantiles, showing that G6PC2 serves as a guardian against low glucose levels without driving dangerous hyperglycemia, which may explain the lack of association with diabetes risk.

bioinformatics↗

Quantile-specific confounding: correction for subtle population stratification via quantile regression

Subtle population structure remains a significant concern in genome-wide association studies. Using human height as an example, we show how quantile regression, a natural extension of linear regression, can better correct for subtle population structure due to its inherent ability to adjust for quantile-specific effects of covariates such as principal components. We utilize data from the UK biobank and the SardiNIA/ProgeNIA project for demonstration.

genetics↗