bioRxiv · 10.1101/388934
ChIP-seq meta-analysis yields high quality training sets for enhancer classification
Abstract
Genome-wide prediction of enhancers depends on high-quality positive and negative training sets. The use of ChIP-seq peaks as positive training data can be problematic due to high degrees of indirectly bound regions, and often poor overlap between experimental conditions.\n\nHere we explore meta-analysis of ChIP-seq data to generate high-quality training data for enhancer modeling. Our method is based on rank aggregation and identifies a core set of directly bound regions per transcription factor, exploiting between five and twenty ChIP-seq data sets per factor. We applied this method to six different transcription factors, namely TP53, REST, SOX2, GRHL2, HIF1A and PPARG. Sequence analysis and modeling of recurrently bound enhancers yielded distinct enhancer features for the different factors, whereby binding sites of REST and TP53 are strongly determined by their motif; binding of GRHL2 and SOX2 is determined by nucleosome positioning; and binding of PPARG and HIF1A depends on other transcription factors. In conclusion, meta-analysis of ChIP-seq peaks, and centering on motifs, allowed discovering new properties of transcription factor binding.
Explore related subjects
Keep this discovery
Imrichova, H., Aerts, S.. 2018-08-09. ChIP-seq meta-analysis yields high quality training sets for enhancer classification. https://doi.org/10.1101/388934
Cite the original work for its findings. Save a collection to share your selection of sources.