bioRxiv Science⌕ Search

Biology subjects

Chundru, K.

Publications and source records attributed to Chundru, K..

2 recordsLinked to original sources

Whole-genome sequencing analysis of anthropometric traits in 672,976 individuals reveals convergence between rare and common genetic associations

Genetic association studies have mostly focussed on common variants from genotyping arrays or rare protein-coding variants from exome sequencing. Here, we used whole-genome sequence (WGS) data in 672,976 individuals of diverse ancestry to evaluate the contribution and architecture of rare non-coding variants to three commonly studied anthropometric traits: height, body mass index (BMI) and waist-hip ratio adjusted for BMI (WHRadjBMI). Analysing 447,461 individuals in UK Biobank for discovery and 225,515 individuals in All of Us for replication, we identified 90 novel rare and low-frequency single variant associations. This includes two independent rare variants upstream of IGF2BP2 that both substantially reduce WHRadjBMI, but have distinct effects on other adiposity traits. We identified 135 coding variant aggregates, several of which were missed by exome sequencing studies. For example, UBR3 protein-truncating variants were associated with a 2.7kg/m2 increase in BMI. We additionally identified 51 non-coding variant aggregate associations, including in the 5UTR of FGF18 (a highly constrained gene with no previously reported coding associations) associated with up to 6cm effects on height. We show that 97% of rare variant associations occur near GWAS loci demonstrating convergence of rare and common variant associations. Finally, we show that ultra rare variants (MAF<0.01%) explain a small fraction of heritability (<10%) compared to common variants for these traits, that heritability is largely shared across ancestries, and that this heritability is concentrated at or near common variant loci. Our work demonstrates the importance of large-scale WGS for fully understanding the genetic architecture of complex traits.

genetics↗

Whole genome sequencing analysis identifies rare, large-effect non-coding variants and regions associated with circulating protein levels

The role of non-coding rare variation in common phenotypes is largely unknown, due to a lack of whole-genome sequence data, and the difficulty of categorising non-coding variants into biologically meaningful regulatory units. To begin addressing these challenges, we performed a cis association analysis using whole-genome sequence data, consisting of 391 million variants and 1,450 circulating protein levels in [~]20,000 UK Biobank participants. We identified 777 independent rare non-coding single variants associated with circulating protein levels (P<1x10-9), after conditioning on protein-coding and common associated variants. Rare non-coding aggregate testing identified 108 conditionally independent regulatory regions. Unlike protein-coding variation, rare non-coding genetic variation was almost as likely to increase as decrease protein levels. The regions we identified overlapped predicted tissue-specific enhancers more than promoters, suggesting they represent tissue-specific regulatory regions. Our results have important implications for the identification, and role, of rare non-coding variation associated with common human phenotypes.

genetics↗