bioRxiv ScienceSearch

Biology subjects

Png, G.

Publications and source records attributed to Png, G..

2 recordsLinked to original sources

Whole genome sequencing analysis of the cardiometabolic proteome

The human proteome is a crucial intermediate between complex diseases and their genetic and environmental components, and an important source of drug development targets and biomarkers. Here, we comprehensively assess the genetic architecture of 257 circulating protein biomarkers of cardiometabolic relevance through high-depth (22.5x) whole-genome sequencing (WGS) in 1,328 individuals. We discover 131 independent sequence variant associations (P<7.45x10-11) across the allele frequency spectrum, all of which replicate in an independent cohort (n=1,605, 18.4x WGS). We identify for the first time replicating evidence for rare-variant cis-acting protein quantitative trait loci for five genes, involving both coding and non-coding variation. We construct and validate polygenic scores that explain up to 45% of protein level variation. We find causal links between protein levels and disease risk, identifying high-value biomarkers and drug development targets.

genetics

Population-wide copy number variation calling using variant call format files from 6,898 individuals

MotivationCopy number variants (CNVs) are large deletions or duplications at least 50 to 200 base pairs long. They play an important role in multiple disorders, but accurate calling of CNVs remains challenging. Most current approaches to CNV detection use raw read alignments, which are computationally intensive to process.\n\nResultsWe use a regression tree-based approach to call CNVs from whole-genome sequencing (WGS, > 18x) variant call-sets in 6,898 samples across four European cohorts, and describe a rich large variation landscape comprising 1,320 CNVs. 61.8% of detected events have been previously reported in the Database of Genomic Variants. 23% of high-quality deletions affect entire genes, and we recapitulate known events such as the GSTM1 and RHD gene deletions. We test for association between the detected deletions and 275 protein levels in 1,457 individuals to assess the potential clinical impact of the detected CNVs. We describe the LD structure and copy number variation underlying the association between levels of the CCL3 protein and a complex structural variant (MAF = 0.15, p = 3.6x10-12) affecting CCL3L3, a paralog of the CCL3 gene. We also identify a cis- association between a low-frequency NOMO1 deletion and the protein product of this gene (MAF = 0.02, p = 2.2x10-7), for which no cis- or trans- single nucleotide variant-driven protein quantitative trait locus (pQTL) has been documented to date. This work demonstrates that existing population-wide WGS call-sets can be mined for CNVs with minimal computational overhead, delivering insight into a less well-studied, yet potentially impactful class of genetic variant.\n\nAvailabilityThe regression tree based approach, UN-CNVc, is available as an R and bash executable on GitHub at https://github.com/agilly/un-cnvc.\n\nContacteleftheria.zeggini@helmholtz-muenchen.de; arthur.gilly@helmholtz-muenchen.de\n\nSupplementary InformationSupplementary information is appended.

bioinformatics