bioRxiv Science⌕ Search

Biology subjects

Modolo, E.

Publications and source records attributed to Modolo, E..

3 recordsLinked to original sources

Discrepancies between ChIP-seq and CUT&Tag histone mark profiles are explained by GC content and chromatin accessibility

Chromatin profiling methods, such as ChIP-seq (chromatin immunoprecipitation followed by sequencing), are used to characterize the genomic localization of DNA-associated proteins. While ChIP and ChIP-seq have been used for decades, an orthogonal approach, CUT&Tag (Cleavage Under Targets and Tagmentation), is gaining popularity as an efficient and cost-effective alternative. Although previous comparative studies note discrepancies in signal-to-noise ratios and detection bias at certain genomic regions, many differences between ChIP-seq and CUT&Tag results remain largely unreconciled. Here, we systematically investigate the disagreeing signals captured by these two methods across well annotated genomic regions. We assess multiple histone mark profiles generated by different groups in two cell lines (K562 and MCF-7). Overall, our analysis indicates that compared to ChIP-seq, CUT&Tag may have limited sensitivity in low-GC environments and, as previously observed, increased signal in hyper-accessible chromatin. Notably, within GC-poor regions of active gene bodies and Polycomb-repressed domains, CUT&Tag exhibits a loss of H3K36me3 and H3K27me3 signal, respectively, where occupancy of these histone marks is otherwise expected. Further, active promoters with discrepant H3K4me3 and H3K27ac signal between the two assays differ systematically in GC content and chromatin accessibility. Promoters differentially enriched for CUT&Tag signal relative to ChIP-seq generally show higher, broader GC-content profiles and higher DNase-seq and ATAC-seq signal, while promoters enriched for ChIP-seq signal harbor the opposite characteristics. A local bias for high GC content and/or chromatin accessibility in CUT&Tag may also explain its differing signal patterns at nucleosome-depleted regions (NDRs) compared to ChIP-seq and MNase-seq promoter profiles. Altogether, our results highlight that studies focused on profiling the intensity and structure of histone modification occupancy can be sensitive to potential biases of CUT&Tag to GC content and chromatin accessibility. These characteristics should be accounted for when selecting a chromatin profiling approach, analyzing and interpreting data as well as drawing biological conclusions.

genomics↗

Multiplexed measurements of protein-protein interactions and protein abundance across cellular conditions using Prod&PQ-seq

Methods to profile protein-protein interactions (PPIs) have limited scalability and can only study a handful of conditions and/or targets. Here, we introduce Prod&PQ-seq, a framework for multiplexed detection and quantification of PPIs and proteins. Our framework uses cross-linked cells, antibody-oligonucleotide conjugates (ab-oligos), and captures PPIs by the DNA-caliper, a specialized oligonucleotide for bidirectional priming of proximal ab-oligos. We benchmarked Prod&PQ-seq using recombinant complexes, titrations and cell mixture experiments and show that our framework is quantitative, reproducible, sensitive and specific. Applying Prod&PQ-seq to study Polycomb Repressive Complex 2 (PRC2) shows that EZH2 inhibition and expression of the oncohistone H3.3K27M weakens both PRC2-H3K27me3 interactions and PPIs within PRC2. Further, H3.1K27M and H3.3K27M variants lead to distinct PPI profiles such as the intensity of H3K27ac-K27M or H3K27ac-EED. Together, Prod&PQ-seq enables detection of changes in PPI composition and intensity and protein quantification across biological conditions, small molecule inhibition and genetic perturbations. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=197 HEIGHT=200 SRC="FIGDIR/small/697286v2_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@1ed0805org.highwire.dtl.DTLVardef@a9c0d9org.highwire.dtl.DTLVardef@b411fborg.highwire.dtl.DTLVardef@8a12e_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

Systematic evaluation of the impact of promoter proximal short tandem repeats on expression

Genetic variation at thousands of short tandem repeats (STRs), which consist of consecutive repeated sequences of 1-6bp, has been statistically associated with gene expression and other molecular phenotypes in humans. However, the causality and regulatory mechanisms for most of these STRs remains unknown. Massively parallel reporter assays (MPRA) enable testing the regulatory activity of a large number of synthesized variants, but have not been applied to STRs due to experimental and computational challenges. Here, we optimized an MPRA framework based on random barcoding to study the impact of variation in repeat copy number on expression. We first performed an MPRA on sequences derived from 30,516 promoter-proximal STR loci along with up to 152bp of genomic context, testing 3-4 variants with differing repeat copy numbers for each locus in HEK293T cells. We identified 1,366 loci with significant associations between repeat copy number and expression, which were enriched for positive effect sizes (P=2.08e-110). We then designed a second MPRA in which we performed deeper perturbations, including systematic manipulation of the repeat unit sequence, orientation, and copy number, with 200-300 perturbations for each of the 300 loci with the strongest signals. Our results revealed that the repeat unit sequence is the primary driver of differences in the relationship between copy number and expression across loci, whereas orientation and flanking sequence have weaker effects, primarily for AT-rich repeat units. The high resolution of these perturbations enabled us to detect non-linear effects, most notably for AAAC/GTTT repeats, which emerge only beyond a certain copy number threshold. Finally, we observed that a subset of STRs in our library show expression levels that are tightly linked with predicted DNA secondary structure formation. We repeated our perturbation MPRA in HeLa S3 cells under wildtype and RNase H1 knockdown conditions, which, via reduction in RNase H1 activity, are expected to hinder resolution of R-loops. This demonstrated that associations between copy number and expression at G-quadruplex-forming CCCCG/CGGGG repeats are particularly sensitive to loss of RNase H1, providing support for an R-loop mediated mechanism for these repeats. Altogether, we establish STRs as a critical component of the non-coding regulatory grammar and provide a framework for understanding how this dynamic form of genetic variation shapes gene expression.

genomics↗