bioRxiv ScienceSearch

Biology subjects

Ebert, P.

Publications and source records attributed to Ebert, P..

7 recordsLinked to original sources

Fast Detection of Differential Chromatin Domains with SCIDDO

The generation of genome-wide maps of histone modifications using chromatin immunoprecipitation sequencing (ChIP-seq) is a common approach to dissect the complexity of the epigenome. However, interpretation and differential analysis of histone ChIP-seq datasets remains challenging due to the genomic co-occurrence of several marks and their difference in genomic spread. Here we present SCIDDO, a fast statistical method for the detection of differential chromatin domains (DCDs) from chromatin state maps. DCD detection simplifies relevant tasks such as the characterization of chromatin changes in differentially expressed genes or the examination of chromatin dynamics at regulatory elements. SCIDDO is available at github.com/ptrebert/sciddo

bioinformatics

Epigenome-based prediction of gene expression across species

BackgroundCross-species studies of epigenetic regulation have great potential, yet most epige-nome mapping has focused on human, mouse, and a small number of other model organisms. Here we explore whether existing reference epigenome collections can be leveraged for analyzing other species, by extrapolation and predictive transfer of epigenome information from established model organisms to less well annotated non-model organisms.\n\nResultsWe developed a methodology for cross-species mapping of epigenome data, which we used for predicting tissue-specific gene expression across twelve mammalian and one avian species. Specifically, we trained gradient boosting classifiers to predict gene expression status from reference epigenome data in human and mouse, and we applied these classifiers to epigenome profiles that were computationally transferred between species. The resulting predictions indeed identified tissue-specific differences in gene expression in the target species, thus providing initial validation of the concept of cross-species epigenome extrapolation.\n\nConclusionsOur study establishes a workflow for cross-species epigenome mapping and epigenome-based prediction of gene expression, highlighting the future potential of using epigenome maps from reference species to annotate a potentially large number of target species.

genomics

Integrative analysis of single cell expression data reveals distinct regulatory states in bidirectional promoters

BackgroundBidirectional promoters (BPs) are prevalent in eukaryotic genomes. However, it is poorly understood how the cell integrates different epigenomic information, such as transcription factor (TF) binding and chromatin marks, to drive gene expression at BPs. Single cell sequencing technologies are revolutionizing the field of genome biology. Therefore, this study focuses on the integration of single cell RNA-seq data with bulk ChIP-seq and other epigenetics data, for which single cell technologies are not yet established, in the context of BPs.\n\nResultsWe performed integrative analyses of novel human single cell RNA-seq (scRNA-seq) data with bulk ChIP-seq and other epigenetics data. scRNA-seq data revealed distinct transcription states of BPs that were previously not recognized. We find associations between these transcription states to distinct patterns in structural gene features, DNA accessibility, histone modification, DNA methylation and TF binding profiles.\n\nConclusionsOur results suggest that a complex interplay of all of these elements is required to achieve BP-specific transcriptional output in this specialized promoter configuration. Further, our study implies that novel statistical methods can be developed to deconvolute masked subpopulations of cells measured with different bulk epigenomic assays using scRNA-seq data.

genomics

Partially methylated domains are hallmarks of a cell specific epigenome topology

BackgroundPartially methylated domains, PMDs, are extended regions in the genome exhibiting a reduced average DNA-methylation level. PMDs cover gene-poor and transcriptionally inactive regions and tend to be heterochromatic. Here, we present a first comprehensive comparative analysis of PMDs across more than 190 WGBS methylomes of human and mouse cells providing a deep insight into structural and functional features associated with PMDs.\n\nResultsPMDs are ubiquitous signatures covering up to 75% of the genome in human and mouse cells irrespective of their tissue or cell origin. Additionally, each cell type comes with a distinct set of specific PMDs, and genes expressed in such PMDs show a strong cell type effect. Demethylation strength varies in PMDs with a tendency towards a more pronounced effect in differentiating and replicating cells. The strongest demethylation is observed in highly proliferating and immortal cancer cell lines. A decrease of DNA-methylation within PMDs tends to be linked to an increase in heterochromatic histone marks and a decrease of gene expressions. Characteristic combinations of heterochromatic signatures in PMDs are linked to domains of early, middle and late DNA-replication.\n\nConclusionPMDs are prominent signatures of long-range epigenomic organization. Integrative analysis identifies PMDs as important general, lineage- and cell-type specific topological features. PMD changes are hallmarks of cell differentiation. Demethylation of PMDs combined with increased heterochromatic marks is a feature linked to enhanced cell proliferation. In combination with broad histone marks PMDs demarcate distinct domains of late DNA-replication.

bioinformatics

Temporal epigenomic profiling identifies AHR as dynamic super-enhancer controlled regulator of mesenchymal multipotency

Temporal data on gene expression and context-specific open chromatin states can improve identification of key transcription factors (TFs) and the gene regulatory networks (GRNs) controlling cellular differentiation. However, their integration remains challenging. Here, we delineate a general approach for data-driven and unbiased identification of key TFs and dynamic GRNs, called EPIC-DREM. We generated time-series transcriptomic and epigenomic profiles during differentiation of mouse multipotent bone marrow stromal cells (MSCs) towards adipocytes and osteoblasts. Using our novel approach we constructed time-resolved GRNs for both lineages. To prioritize the identified shared regulators, we mapped dynamic super-enhancers in both lineages and associated them to target genes with correlated expression profiles. We identified aryl hydrocarbon receptor (AHR) and Glis family zinc finger 1 (GLIS1) as mesenchymal key TFs controlled by dynamic MSC-specific super-enhancers that become repressed in both lineages. AHR and GLIS1 control differentiation-induced genes and we propose they function as guardians of mesenchymal multipotency.

bioinformatics

All Fingers Are Not The Same: Handling Variable-Length Sequences In A Discriminative Setting Using Conformal Multi-Instance Kernels

Most string kernels for comparison of genomic sequences are generally tied to using (absolute) positional information of the features in the individual sequences. This poses limitations when comparing variable-length sequences using such string kernels. For example, profiling chromatin interactions by 3C-based experiments results in variable-length genomic sequences (restriction fragments). Here, exact position-wise occurrence of signals in sequences may not be as important as in the scenario of analysis of the promoter sequences, that typically have a transcription start site as reference. Existing position-aware string kernels have been shown to be useful for the latter scenario.\n\nIn this work, we propose a novel approach for sequence comparison that enables larger positional freedom than most of the existing approaches, can identify a possibly dispersed set of features in comparing variable-length sequences, and can handle both the aforementioned scenarios. Our approach, CoMIK, identifies not just the features useful towards classification but also their locations in the variable-length sequences, as evidenced by the results of three binary classification experiments, aided by recently introduced visualization techniques. Furthermore, we show that we are able to efficiently retrieve and interpret the weight vector for the complex setting of multiple multi-instance kernels.

bioinformatics

Combining transcription factor binding affinities with open-chromatin data for accurate gene expression prediction

The binding and contribution of transcription factors (TF) to cell specific gene expression is often deduced from open-chromatin measurements to avoid costly TF ChIP-seq assays. Thus, it is important to develop computational methods for accurate TF binding prediction in open-chromatin regions (OCRs). Here, we report a novel segmentation-based method, TEPIC, to predict TF binding by combining sets of OCRs with position weight matrices. TEPIC can be applied to various open-chromatin data, e.g. DNaseI-seq and NOMe-seq. Additionally, Histone-Marks (HMs) can be used to identify candidate TF binding sites. TEPIC computes TF affinities and uses open-chromatin/HM signal intensity as quantitative measures of TF binding strength. Using machine learning, we find low affinity binding sites to improve our ability to explain gene expression variability compared to the standard presence/absence classification of binding sites. Further, we show that both footprints and peaks capture essential TF binding events and lead to a good prediction performance. In our application, gene-based scores computed by TEPIC with one open-chromatin assay nearly reach the quality of several TF ChIP-seq datasets. Finally, these scores correctly predict known transcriptional regulators as illustrated by the application to novel DNaseI-seq and NOMe-seq data for primary human hepatocytes and CD4+ T-cells, respectively.

bioinformatics