bioRxiv Science⌕ Search

Biology subjects

James, D. Q.

Publications and source records attributed to James, D. Q..

2 recordsLinked to original sources

Optimized ChIP-exo for mammalian cells and patterned sequencing flow cells

By combining chromatin immunoprecipitation (ChIP) with an exonuclease digestion of protein-bound DNA fragments, ChIP-exo characterizes genome-wide protein-DNA interactions at near base-pair resolution. However, the widespread adoption of ChIP-exo has been hindered by several technical challenges, including lengthy protocols, the need for multiple custom reactions, and incompatibilities with recent Illumina sequencing platforms. To address these barriers, we systematically optimized and adapted the ChIP-exo library construction protocol for the unique requirements of mammalian cells and current sequencing technologies. We introduce a Mammalian-Optimized ChIP-exo (MO-ChIP-exo) protocol that builds upon previous ChIP-exo protocols with systematic optimization of crosslinking, harvesting, and library construction. We validate MO-ChIP-exo by comparing it to previously published ChIP-exo protocols and demonstrate its adaptability to both suspension (K562) and adherent (HepG2, mESC) cell lines. This improved protocol provides a more robust and efficient method for generating high-quality ChIP-exo libraries from mammalian cells. SUMMARYChIP-exo is a genome-wide protein-DNA binding assay with unrivalled resolution, but its widespread adoption has been hindered by technical challenges, particularly when applied to mammalian cells or when used with recent sequencing platforms. We introduce a Mammalian-Optimized ChIP-exo (MO-ChIP-exo) protocol with key modifications that overcome previous technical hurdles. We demonstrate that our optimized protocol produces high-quality data comparable to previously published protocols and is adaptable for use with both suspension (K562) and adherent (HepG2, mESC) cell lines.

genomics↗

Allo: Accurate allocation of multi-mapped reads enables regulatory element analysis at repeats

Transposable elements (TEs) and other repetitive regions have been shown to contain gene regulatory elements, including transcription factor binding sites. Unfortunately, regulatory elements harbored by repeats have proven difficult to characterize using short-read sequencing assays such as ChIP-seq or ATAC-seq. Most regulatory genomics analysis pipelines discard "multi-mapped" reads that align equally well to multiple genomic locations. Since multi-mapped reads arise predominantly from repeats, current analysis pipelines fail to detect a substantial portion of regulatory events that occur in repetitive regions. To address this shortcoming, we developed Allo, a new approach to allocate multi-mapped reads in an efficient, accurate, and user-friendly manner. Allo combines probabilistic mapping of multi-mapped reads with a convolutional neural network that recognizes the read distribution features of potential peaks, offering enhanced accuracy in multi-mapping read assignment. Allo also provides read-level output in the form of a corrected alignment file, making it compatible with existing regulatory genomics analysis pipelines and downstream peak-finders. In a demonstration application on CTCF ChIP-seq data, we show that Allo results in the discovery of thousands of new CTCF peaks. Many of these peaks contain the expected cognate motif and/or serve as TAD boundaries. We additionally apply Allo to a diverse collection of ENCODE ChIP-seq datasets, resulting in multiple previously unidentified interactions between transcription factors and repetitive element families. Finally, we show that Allo may be particularly effective in identifying ChIP-seq peaks in younger TEs, which hold evolutionary significance due to their emergence during human evolution from primates.

bioinformatics↗