bioRxiv ScienceSearch

Biology subjects

Delaneau, O.

Publications and source records attributed to Delaneau, O..

7 recordsLinked to original sources

Expression estimation and eQTL mapping for HLA genes with a personalized pipeline

The HLA (Human Leukocyte Antigens) genes are well-documented targets of balancing selection, and variation at these loci is associated with many disease phenotypes. Variation in expression levels also influences disease susceptibility and resistance, but little information exists about the regulation and population-level patterns of expression due to the difficulty in mapping short reads to these highly polymorphic loci, and in accounting for the existence of several paralogues. We developed a computational pipeline to accurately estimate expression for HLA genes based on RNA-seq, improving both locus-level and allele-level estimates. First, reads are aligned to all known HLA sequences in order to infer HLA genotypes, then quantification of expression is carried out using a personalized index. We use simulations to show that expression estimates are not biased due to divergence from the reference genome. We applied our pipeline to GEUVADIS dataset, and compared the quantifications to those obtained with reference transcriptome, and found that a substantial portion of the variation captured by the HLA-personalized index in not captured by the standard index (23%). We describe the impact of the HLA-personalized approach on downstream analyses for seven HLA loci (HLA-A, HLA-B, HLA-C, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRB1). Although the influence of the HLA-personalized approach is modest for eQTL mapping, the p-values and the causality of the eQTLs obtained are better than when the reference transcriptome is used. Finally, we integrate information on HLA-allele level expression with the eQTL findings to show that the HLA allele is an important layer of variation to understand HLA regulation.

genomics

A genome-wide association study for shared risk across major psychiatric disorders in a nation-wide birth cohort implicates fetal neurodevelopment as a key mediator

There is mounting evidence that seemingly diverse psychiatric disorders share genetic etiology, but the biological substrates mediating this overlap are not well characterized. Here, we leverage the unique iPSYCH study, a nationally representative cohort ascertained through clinical psychiatric diagnoses indicated in Danish national health registers. We confirm previous reports of individual and cross-disorder SNP-heritability for major psychiatric disorders and perform a cross-disorder genome-wide association study. We identify four novel genome-wide significant loci encompassing variants predicted to regulate genes expressed in radial glia and interneurons in the developing neocortex during midgestation. This epoch is supported by partitioning cross-disorder SNP-heritability which is enriched at regulatory chromatin active during fetal neurodevelopment. These findings indicate that dysregulation of genes that direct neurodevelopment by common genetic variants results in general liability for many later psychiatric outcomes.

genetics

Germline determinants of the somatic mutation landscape in 2,642 cancer genomes

Cancers develop through somatic mutagenesis, however germline genetic variation can markedly contribute to tumorigenesis via diverse mechanisms. We discovered and phased 88 million germline single nucleotide variants, short insertions/deletions, and large structural variants in whole genomes from 2,642 cancer patients, and employed this genomic resource to study genetic determinants of somatic mutagenesis across 39 cancer types. Our analyses implicate damaging germline variants in a variety of cancer predisposition and DNA damage response genes with specific somatic mutation patterns. Mutations in the MBD4 DNA glycosylase gene showed association with elevated C>T mutagenesis at CpG dinucleotides, a ubiquitous mutational process acting across tissues. Analysis of somatic structural variation exposed complex rearrangement patterns, involving cycles of templated insertions and tandem duplications, in BRCA1-deficient tumours. Genome-wide association analysis implicated common genetic variation at the APOBEC3 gene cluster with reduced basal levels of somatic mutagenesis attributable to APOBEC cytidine deaminases across cancer types. We further inferred over a hundred polymorphic L1/LINE elements with somatic retrotransposition activity in cancer. Our study highlights the major impact of rare and common germline variants on mutational landscapes in cancer.

genomics

Hundreds of putative non-coding cis-regulatory drivers in chronic lymphocytic leukaemia and skin cancer

Perturbations of the coding genome and their role in cancer development have been studied extensively. However, the non-coding genomes contribution in cancer is poorly understood (1), not only because it is difficult to define the non-coding regulatory regions and the genes they regulate, but also because there is limited power owing to the regulatory regions small size. In this study, we try to resolve this issue by defining modules of coordinated non-coding regulatory regions of genes (Cis Regulatory Domains or CRDs). To do so, we use the correlation between histone modifications, assayed by ChIP-seq, in population samples of immortalized B-cells and skin fibroblasts. We screen for CRDs that accumulate an excess of somatic mutations in chronic lymphocytic leukaemia (CLL) and skin cancer, which affect these cell types, after accounting for somatic mutational patterns and biases. At 5% FDR, we find 90 CRDs with significant excess somatic of mutations in CLL, 60 of which regulate 126 genes, and in skin cancer 59 significant CRDs, 25 of which regulate 37 genes. The genes these CRDs regulate include ones already implicated in tumorigenesis, and are enriched in pathways already implicated in the respective cancers, like the B-cell receptor signalling pathway in CLL and the TGF{beta} signalling pathway in skin cancer. We discover that the somatic mutations in the significant CRDs of CLL are hitting bases more likely to be functional than the mutations in non-significant CRDs. Moreover, in both cancers, mutational signatures observed in the regulatory regions of significant CRDs deviate significantly from their null sequences. Both results indicate selection acting on CRDs during tumorigenesis. Finally, we find that the transcription factor biding sites that are disturbed by the somatic mutations in significant CRDs are enriched for factors known to be involved in cancer development. We are describing a new powerful approach to discover non-coding regions involved in tumorigenesis in CLL and skin cancer and this approach could be generalized to other cancers.

cancer biology

Intra- and inter-chromosomal chromatin interactions mediate genetic effects on regulatory networks

Genome-wide studies on the genetic basis of gene expression and the structural properties of chromatin have considerably advanced our understanding of the function of the human genome. However, it remains unclear how structure relates to function and, in this work, we aim at bridging both by assembling a dataset that combines the activity of regulatory elements (e.g. enhancers and promoters), expression of genes and genetic variations of 317 individuals and across two cell types. We show that the regulatory activity is structured within 12,583 Cis Regulatory Domains (CRDs) that are cell type specific and highly reflective of the local (i.e. Topologically Associating Domains) and global (i.e. A/B nuclear compartments) nuclear organization of the chromatin. These CRDs essentially delimit the sets of active regulatory elements involved in the transcription of most genes, thereby capturing complex regulatory networks in which the effects of regulatory variants are propagated and combined to finally mediate expression Quantitative Trait Loci. Overall, our analysis reveals the complexity and specificity of cis and trans regulatory networks and their perturbation by genetic variation.

genomics

Genome-wide genetic data on ~500,000 UK Biobank participants

The UK Biobank project is a large prospective cohort study of ~500,000 individuals from across the United Kingdom, aged between 40-69 at recruitment. A rich variety of phenotypic and health-related information is available on each participant, making the resource unprecedented in its size and scope. Here we describe the genome-wide genotype data (~805,000 markers) collected on all individuals in the cohort and its quality control procedures. Genotype data on this scale offers novel opportunities for assessing quality issues, although the wide range of ancestries of the individuals in the cohort also creates particular challenges. We also conducted a set of analyses that reveal properties of the genetic data - such as population structure and relatedness - that can be important for downstream analyses. In addition, we phased and imputed genotypes into the dataset, using computationally efficient methods combined with the Haplotype Reference Consortium (HRC) and UK10K haplotype resource. This increases the number of testable variants by over 100-fold to ~96 million variants. We also imputed classical allelic variation at 11 human leukocyte antigen (HLA) genes, and as a quality control check of this imputation, we replicate signals of known associations between HLA alleles and many common diseases. We describe tools that allow efficient genome-wide association studies (GWAS) of multiple traits and fast phenome-wide association studies (PheWAS), which work together with a new compressed file format that has been used to distribute the dataset. As a further check of the genotyped and imputed datasets, we performed a test-case genome-wide association scan on a well-studied human trait, standing height.

genetics

Predicting causal variants affecting expression using whole genome sequence and RNA-seq from multiple human tissues.

Genetic association mapping produces statistical links between phenotypes and genomic regions, but identifying the causal variants themselves remains difficult. Complete knowledge of all genetic variants, as provided by whole genome sequence (WGS), will help, but is currently financially prohibitive for well powered GWAS studies. To explore the advantages of WGS in a well powered setting, we performed eQTL mapping using WGS and RNA-seq, and showed that the lead eQTL variants called using WGS are more likely to be causal. We derived properties of the causal variant from simulation studies, and used these to propose a method for implicating likely causal SNPs. This method predicts that 25% - 70% of the causal variants lie in open chromatin regions, depending on tissue and experiment. Finally, we identify a set of high confidence causal variants and show that they are more enriched in GWAS associations than other eQTL. Of these, we find 65 associations with GWAS traits and show examples where the gene implicated by expression has been functionally validated as relevant for complex traits.

genomics