bioRxiv ScienceSearch

Biology subjects

Patrick Deelen

Publications and source records attributed to Patrick Deelen.

5 recordsLinked to original sources

A high-quality reference panel reveals the complexity and distribution of structural genome changes in a human population

Structural variation (SV) represents a major source of differences between individual human genomes and has been linked to disease phenotypes. However, the majority of studies provide neither a global view of the full spectrum of these variants nor integrate them into reference panels of genetic variation.\n\nHere, we analyse whole genome sequencing data of 769 individuals from 250 Dutch families, and provide a haplotype-resolved map of 1.9 million genome variants across 9 different variant classes, including novel forms of complex indels, and retrotransposition-mediated insertions of mobile elements and processed RNAs. A large proportion are previously under reported variants sized between 21 and 100bp. We detect 4 megabases of novel sequence, encoding 11 new transcripts. Finally, we show 191 known, trait-associated SNPs to be in strong linkage disequilibrium with SVs and demonstrate that our panel facilitates accurate imputation of SVs in unrelated individuals. Our findings are essential for genome-wide association studies.

Genetics

Disease variants alter transcription factor levels and methylation of their binding sites

Most disease associated genetic risk factors are non-coding, making it challenging to design experiments to understand their functional consequences1,2. Identification of expression quantitative trait loci (eQTLs) has been a powerful approach to infer downstream effects of disease variants but the large majority remains unexplained.3,4. The analysis of DNA methylation, a key component of the epigenome5, offers highly complementary data on the regulatory potential of genomic regions6,7. However, a large-scale, combined analysis of methylome and transcriptome data to infer downstream effects of disease variants is lacking. Here, we show that disease variants have wide-spread effects on DNA methylation in trans that likely reflect the downstream effects on binding sites of cis-regulated transcription factors. Using data on 3,841 Dutch samples, we detected 272,037 independent cis-meQTLs (FDR < 0.05) and identified 1,907 trait-associated SNPs that affect methylation levels of 10,141 different CpG sites in trans (FDR < 0.05), an eight-fold increase in the number of downstream effects that was known from trans-eQTL studies3,8,9. Trans-meQTL CpG sites are enriched for active regulatory regions, being correlated with gene expression and overlap with Hi-C determined interchromosomal contacts10,11. We detected many trans-meQTL SNPs that affect expression levels of nearby transcription factors (including NFKB1, CTCF and NKX2-3), while the corresponding trans-meQTL CpG sites frequently coincide with its respective binding site. Trans-meQTL mapping therefore provides a strategy for identifying and better understanding downstream functional effects of many disease-associated variants.

Genomics

Hypothesis-free identification of modulators of genetic risk factors

Genetic risk factors often localize in non-coding regions of the genome with unknown effects on disease etiology. Expression quantitative trait loci (eQTLs) help to explain the regulatory mechanisms underlying the association of genetic risk factors with disease. More mechanistic insights can be derived from knowledge of the context, such as cell type or the activity of signaling pathways, influencing the nature and strength of eQTLs. Here, we generated peripheral blood RNA-seq data from 2,116 unrelated Dutch individuals and systematically identified these context-dependent eQTLs using a hypothesis-free strategy that does not require prior knowledge on the identity of the modifiers. Out of the 23,060 significant cis-regulated genes (false discovery rate < 0.05), 2,743 genes (12%) show context-dependent eQTL effects. The majority of those were influenced by cell type composition, revealing eQTLs that are particularly strong in cell types such as CD4+ T-cells, erythrocytes, and even lowly abundant eosinophils. A set of 145 cis-eQTLs were influenced by the activity of the type I interferon signaling pathway and we identified several cis-eQTLs that are modulated by specific transcription factors that bind to the eQTL SNPs. This demonstrates that large-scale eQTL studies in unchallenged individuals can complement perturbation experiments to gain better insight in regulatory networks and their stimuli.

Genetics

An introduction to LifeLines DEEP: study design and baseline characteristics

There is a critical need for population-based prospective cohort studies because they follow individuals before the onset of disease, allowing for studies that can identify biomarkers and disease-modifying effects and thereby contributing to systems epidemiology. This paper describes the design and baseline characteristics of an intensively examined subpopulation of the LifeLines cohort in the Netherlands. For this unique sub-cohort, LifeLines DEEP, additional blood (n=1387), exhaled air (n=1425), fecal samples (n=1248) and gastrointestinal health questionnaires (n=1176) were collected for analysis of the genome, epigenome, transcriptome, microbiome, metabolome and other biological levels. Here, we provide an overview of the different data layers in LifeLines DEEP and present baseline characteristics of the study population including food intake and quality of life. We also describe how the LifeLines DEEP cohort allows for the detailed investigation of genetic, genomic and metabolic variation on a wealth of phenotypic outcomes. Finally, we examine the determinants of gastrointestinal health, an area of particular interest to us that can be addressed by LifeLines DEEP.

Genomics

Calling genotypes from public RNA-sequencing data enables identification of genetic variants that affect gene-expression levels

Given increasing numbers of RNA-seq samples in the public domain, we studied to what extent expression quantitative trait loci (eQTLs) and allele-specific expression (ASE) can be identified in public RNA-seq data while also deriving the genotypes from the RNA-seq reads. 4,978 human RNA-seq runs, representing many different tissues and cell-types, passed quality control. Even though this data originated from many different laboratories, samples reflecting the same cell-type clustered together, suggesting that technical biases due to different sequencing protocols were limited. We derived genotypes from the RNA-seq reads and imputed non-coding variants. In a joint analysis on 1,262 samples combined, we identified cis-eQTLs effects for 8,034 unique genes. Additionally, we observed strong ASE effects for 34 rare pathogenic variants, corroborating previously observed effects on the corresponding protein levels. Given the exponential growth of the number of publicly available RNA-seq samples, we expect this approach will become relevant for studying tissue-specific effects of rare pathogenic genetic variants.

Genetics