bioRxiv ScienceSearch

Biology subjects

Ay, F.

Publications and source records attributed to Ay, F..

10 recordsLinked to original sources

TET enzymes augment AID expression via 5hmC modifications at the Aicda superenhancer

TET enzymes are dioxygenases that promote DNA demethylation by oxidizing the methyl group of 5-methylcytosine (5mC) to 5-hydroxymethylcytosine (5hmC). Here we report a close correspondence between 5hmC-marked regions, chromatin accessibility and enhancer activity in B cells, and a strong enrichment for consensus binding motifs for basic region-leucine zipper (bZIP) transcription factors at TET-responsive genomic regions. Functionally, Tet2 and Tet3 regulate class switch recombination (CSR) in murine B cells by enhancing expression of Aicda, encoding the cytidine deaminase AID essential for CSR. TET enzymes deposit 5hmC, demethylate and maintain chromatin accessibility at two TET-responsive elements, TetE1 and TetE2, located within a superenhancer in the Aicda locus. Transcriptional profiling identified BATF as the bZIP transcription factor involved in TET-dependent Aicda expression. 5hmC is not deposited at TetE1 in activated Batf-deficient B cells, indicating that BATF recruits TET proteins to the Aicda enhancer. Our data emphasize the importance of TET enzymes for bolstering AID expression, and highlight 5hmC as an epigenetic mark that captures enhancer dynamics during cell activation.

immunology

FitHiChIP: Identification of significant chromatin contacts from HiChIP data

Here we describe FitHiChIP (github.com/ay-lab/FitHiChIP), a computational method for identifying chromatin contacts among regulatory regions such as en-hancers and promoters from HiChIP/PLAC-seq data. FitHiChIP jointly models the non-uniform coverage and genomic distance scaling of HiChIP data, captures previously validated enhancer interactions for several genes including MYC and TP53, and recovers contacts genome-wide that are supported by ChIA-PET, pro-moter capture Hi-C and Hi-C data. FitHiChIP also provides a framework for differential contact analysis as showcased in a comparison of HiChIP data we have generated for two distinct immune cell types.

bioinformatics

mHi-C: robust leveraging of multi-mapping reads in Hi-C analysis

Current Hi-C analysis approaches are unable to account for reads that align to multiple locations, and hence underestimate biological signal from repetitive regions of genomes. We developed mHi-C, a multi-read mapping strategy to probabilistically allocate Hi-C multi-reads. mHi-C exhibited superior performance over utilizing only uni-reads and heuristic approaches aimed at rescuing multi-reads on benchmarks. Specifically, mHi-C increased the sequencing depth by an average of 20% leading to higher reproducibility of contact matrices and larger number of significant interactions across biological replicates. The impact of the multi-reads on the identification of novel significant interactions is influenced marginally by relative contribution of multi-reads to the sequencing depth compared to uni-reads, cis-to-trans ratio of contacts, and the broad data quality as reflected by the proportion of mappable reads of datasets. Computational experiments highlighted that in Hi-C studies with short read lengths, mHi-C rescued multi-reads can emulate the effect of longer reads. mHi-C also revealed biologically supported bona fide promoter-enhancer interactions and topologically associating domains involving repetitive genomic regions, thereby unlocking a previously masked portion of the genome for conformation capture studies.

genomics

Identification of cis elements for spatio-temporal control of DNA replication

The temporal order of DNA replication (replication timing, RT) is highly coupled with genome architecture, but cis-elements regulating spatio-temporal control of replication have remained elusive. We performed an extensive series of CRISPR mediated deletions and inversions and high-resolution capture Hi-C of a pluripotency associated domain (DppA2/4) in mouse embryonic stem cells. Whereas CTCF mediated loops and chromatin domain boundaries were dispensable, deletion of three intra-domain prominent CTCF-independent 3D contact sites caused a domain-wide delay in RT, shift in sub-nuclear chromatin compartment and loss of transcriptional activity, These \"early replication control elements\" (ERCEs) display prominent chromatin features resembling enhancers/promoters and individual and pair-wise deletions of the ERCEs confirmed their partial redundancy and interdependency in controlling domain-wide RT and transcription. Our results demonstrate that discrete cis-regulatory elements mediate domain-wide RT, chromatin compartmentalization, and transcription, representing a major advance in dissecting the relationship between genome structure and function.\n\nHighlightsO_LIcis-elements (ERCEs) regulate large scale chromosome structure and function\nC_LIO_LIMultiple ERCEs cooperatively control domain-wide replication\nC_LIO_LIERCEs harbor prominent active chromatin features and form CTCF-independent loops\nC_LIO_LIERCEs enable genetic dissection of large-scale chromosome structure-function.\nC_LI

molecular biology

ExTraMapper: Exon- and Transcript-level mappings for orthologous gene pairs

Access to large-scale genomics and transcriptomics data from various tissues and cell lines allowed the discovery of wide-spread alternative splicing events and alternative promoter usage in mammalians. However, evolutionary studies primarily focus on gene-level orthology relationships, which hinders the importance of transcript-level diversity. Between human and mouse, gene-level orthology is currently present for nearly 16k protein-coding genes spanning a diverse repertoire of over 200k total transcript isoforms. Here we describe a novel method, ExTraMapper, which leverages sequence conservation between exons of a pair of organisms and identifies a fine-scale orthology mapping at the exon and then transcript level. ExTraMapper identifies more than 250k exon, as well as 30k transcript mappings between human and mouse using only sequence and gene annotation information. We demonstrate that ExTraMapper identifies a larger number of exon and transcript mappings compared to previous methods. Further, it identifies exon fusions, splits, and losses due to splice site mutations, and finds mappings between microexons that are previously missed. By reanalysis of RNA-seq data from 13 matched human and mouse tissues, we show that ExTraMapper improves the correlation of transcript-specific expression levels suggesting a more accurate mapping of human and mouse transcripts. ExTraMapper also reports better transcript-level mappings compared to Ensembl orthology for the human proto-oncogene BRAF and its mouse ortholog as well as several other example genes with important isoform-specific functions. ExTraMapper is applicable to any pair of organisms that have orthologous gene pairs and is available at https://github.com/ay-lab/ExTraMapper and http://ay-lab-tools.lji.org/extramapper

bioinformatics

Changes in genome organization of parasite-specific gene families during the Plasmodium transmission stages

The development of malaria parasites throughout their various life cycle stages is controlled by coordinated changes in gene expression. We previously showed that the three-dimensional organization of the P. falciparum genome is strongly associated with gene expression during its replication cycle inside red blood cells. Here, we analyzed genome organization in the P. falciparum and P. vivax transmission stages. Major changes occurred in the localization and interactions of genes involved in pathogenesis and immune evasion, erythrocyte and liver cell invasion, sexual differentiation and master regulation of gene expression. In addition, we observed reorganization of subtelomeric heterochromatin around genes involved in host cell remodeling. Depletion of heterochromatin protein 1 (PfHP1) resulted in loss of interactions between virulence genes, confirming that PfHP1 is essential for maintenance of the repressive center. Overall, our results suggest that the three-dimensional genome structure is strongly connected with transcriptional activity of specific gene families throughout the life cycle of human malaria parasites.

systems biology

Measuring the reproducibility and quality of Hi-C data

Hi-C is currently the most widely used assay to investigate the 3D organization of the genome and to study its role in gene regulation, DNA replication, and disease. However, Hi-C experiments are costly to perform and involve multiple complex experimental steps; thus, accurate methods for measuring the quality and reproducibility of Hi-C data are essential to determine whether the output should be used further in a study. Using real and simulated data, we profile the performance of several recently proposed methods for assessing reproducibility of population Hi-C data, including HiCRep, GenomeDISCO, HiC-Spector and QuASAR-Rep. By explicitly controlling noise and sparsity through simulations, we demonstrate the deficiencies of performing simple correlation analysis on pairs of matrices, and we show that methods developed specifically for Hi-C data produce better measures of reproducibility. We also show how to use established (e.g., ratio of intra to interchromosomal interactions) and novel (e.g., QuASAR-QC) measures to identify low quality experiments. In this work, we assess reproducibility and quality measures by varying sequencing depth, resolution and noise levels in Hi-C data from 13 cell lines, with two biological replicates each, as well as 176 simulated matrices. Through this extensive validation and benchmarking of Hi-C data, we describe best practices for reproducibility and quality assessment of Hi-C experiments. We make all software publicly available at http://github.com/kundajelab/3DChromatin_ReplicateQC to facilitate adoption in the community.

genomics

Using DNase Hi-C techniques to map global and local three-dimensional genome architecture at high resolution

The folding and three-dimensional (3D) organization of chromatin in the nucleus critically impacts genome function. The past decade has witnessed rapid advances in genomic tools for delineating 3D genome architecture. Among them, chromosome conformation capture (3C)-based methods such as Hi-C are the most widely used techniques for mapping chromatin interactions. However, traditional Hi-C protocols rely on restriction enzymes (REs) to fragment chromatin and are therefore limited in resolution. We recently developed DNase Hi-C for mapping 3D genome organization, which uses DNase I for chromatin fragmentation. DNase Hi-C overcomes RE-related limitations associated with traditional Hi-C methods, leading to improved methodological resolution. Furthermore, combining this method with DNA capture technology provides a high-throughput approach (targeted DNase Hi-C) that allows for mapping fine-scale chromatin architecture at exceptionally high resolution. Hence, targeted DNase Hi-C will be valuable for delineating the physical landscapes of cis-regulatory networks that control gene expression and for characterizing phenotype-associated chromatin 3D signatures. Here, we provide a detailed description of method design and step-by-step working protocols for these two methods.\n\nHighlightsO_LIDNase Hi-C, a method for comprehensive mapping of chromatin contacts on a whole-genome scale, is based on random chromatin fragmentation by DNase I digestion instead of sequence-specific restriction enzyme (RE) digestion.\nC_LIO_LITargeted DNase Hi-C, which combines DNase Hi-C with DNA capture technology, is a high-throughput method for mapping fine-scale chromatin architecture of genomic loci of interest at a resolution comparable to that of genomic annotations of functional elements.\nC_LIO_LIDNase Hi-C and targeted DNase Hi-C provide the first high-throughput way to overcome the RE-digestion-associated resolution limit of 3C-based methods.\nC_LIO_LIStep-by-step whole-genome and targeted DNase Hi-C protocols for mapping global and local 3D genome architecture, respectively, are described.\nC_LI

genomics

Identification of copy number variations and translocations in cancer cells from Hi-C data

MotivationEukaryotic chromosomes adapt a complex and highly dynamic three-dimensional (3D) structure, which profoundly affects different cellular functions and outcomes including changes in epigenetic landscape and in gene expression. Making the scenario even more complex, cancer cells harbor chromosomal abnormalities (e.g., copy number variations (CNVs) and translocations) altering their genomes both at the sequence level and at the level of 3D organization. High-throughput chromosome conformation capture techniques (e.g., Hi-C), which are originally developed for decoding the 3D structure of the chromatin, provide a great opportunity to simultaneously identify the locations of genomic rearrangements and to investigate the 3D genome organization in cancer cells. Even though Hi-C data has been used for validating known rearrangements, computational methods that can distinguish rearrangement signals from the inherent biases of Hi-C data and from the actual 3D conformation of chromatin, and can precisely detect rearrangement locations de novo have been missing.\n\nResultsIn this work, we characterize how intra and inter-chromosomal Hi-C contacts are distributed for normal and rearranged chromosomes to devise a new set of algorithms (i) to identify genomic segments that correspond to CNV regions such as amplifications and deletions (HiCnv), (ii) to call inter-chromosomal translocations and their boundaries (HiCtrans) from Hi-C experiments, and (iii) to simulate Hi-C data from genomes with desired rearrangements and abnormalities (AveSim) in order to select optimal parameters for and to benchmark the accuracy of our methods. Our results on 10 different cancer cell lines with Hi-C data show that we identify a total number of 105 amplifications and 45 deletions together with 90 translocations, whereas we identify virtually no such events for two karyotypically normal cell lines. Our CNV predictions correlate very well with whole genome sequencing (WGS) data among chromosomes with CNV events for a breast cancer cell line (r=0.89) and capture most of the CNVs we simulate using Avesim. For HiCtrans predictions, we report evidence from the literature for 30 out of 90 translocations for eight of our cancer cell lines. Further-more, we show that our tools identify and correctly classify relatively understudied rearrangements such as double minutes (DMs) and homogeneously staining regions (HSRs).\n\nConclusionsConsidering the inherent limitations of existing techniques for karyotyping (i.e., missing balanced rearrangements and those near repetitive regions), the accurate identification of CNVs and translocations in a cost-effective and high-throughput setting is still a challenge. Our results show that the set of tools we develop effectively utilize moderately sequenced Hi-C libraries (100-300 million reads) to identify known and de novo chromosomal rearrangements/abnormalities in well-established cancer cell lines. With the decrease in required number of cells and the increase in attainable resolution, we believe that our framework will pave the way towards comprehensive mapping of genomic rearrangements in primary cells from cancer patients using Hi-C.\n\nAvailabilityO_LICNV calling: https://github.com/ay-lab/HiCnv\nC_LIO_LITranslocation calling: https://github.com/ay-lab/HiCtrans\nC_LIO_LIHi-C simulation: https://github.com/ay-lab/AveSim\nC_LI

bioinformatics

An Integrative Framework For Detecting Structural Variations In Cancer Genomes

Structural variants can contribute to oncogenesis through a variety of mechanisms, yet, despite their importance, the identification of structural variants in cancer genomes remains challenging. Here, we present an integrative framework for comprehensively identifying structural variation in cancer genomes. For the first time, we apply next-generation optical mapping, high-throughput chromosome conformation capture (Hi-C), and whole genome sequencing to systematically detect SVs in a variety of cancer cells.\n\nUsing this approach, we identify and characterize structural variants in up to 29 commonly used normal and cancer cell lines. We find that each method has unique strengths in identifying different classes of structural variants and at different scales, suggesting that integrative approaches are likely the only way to comprehensively identify structural variants in the genome. Studying the impact of the structural variants in cancer cell lines, we identify widespread structural variation events affecting the functions of non-coding sequences in the genome, including the deletion of distal regulatory sequences, alteration of DNA replication timing, and the creation of novel 3D chromatin structural domains.\n\nThese results underscore the importance of comprehensive structural variant identification and indicate that non-coding structural variation may be an underappreciated mutational process in cancer genomes.

genomics