bioRxiv ScienceSearch

Biology subjects

Greenleaf, W. J.

Publications and source records attributed to Greenleaf, W. J..

17 recordsLinked to original sources

Coupled single-cell CRISPR screening and epigenomic profiling reveals causal gene regulatory networks

Here we present Perturb-ATAC, a method which combines multiplexed CRISPR interference or knockout with genome-wide chromatin accessibility profiling in single cells, based on the simultaneous detection of CRISPR guide RNAs and open chromatin sites by assay of transposase-accessible chromatin with sequencing (ATAC-seq). We applied Perturb-ATAC to transcription factors (TFs), chromatin-modifying factors, and noncoding RNAs (ncRNAs) in [~]4,300 single cells, encompassing more than 63 unique genotype-phenotype relationships. Perturb-ATAC in human B lymphocytes uncovered regulators of chromatin accessibility, TF occupancy, and nucleosome positioning, and identified a hierarchical organization of TFs that govern B cell state, variation, and disease-associated cis-regulatory elements. Perturb-ATAC in primary human epidermal cells revealed three sequential modules of cis-elements that specify keratinocyte fate, orchestrated by the TFs JUNB, KLF4, ZNF750, CEBPA, and EHF. Combinatorial deletion of all pairs of these TFs uncovered their epistatic relationships and highlighted genomic co-localization as a basis for synergistic interactions. Thus, Perturb-ATAC is a powerful and general strategy to dissect gene regulatory networks in development and disease.\n\nHighlightsO_LIA new method for simultaneous measurement of CRISPR perturbations and chromatin state in single cells.\nC_LIO_LIPerturb-ATAC reveals regulatory factors that control cis-element accessibility, trans-factor occupancy, and nucleosome positioning.\nC_LIO_LIPerturb-ATAC reveals regulatory modules of coordinated trans-factor activity in B lymphoblasts.\nC_LIO_LIKeratinocyte differentiation is orchestrated by synergistic activities of co-binding TFs on cis-elements.\nC_LI

genomics

Landscape of stimulation-responsive chromatin across diverse human immune cells

The immune system is controlled by a balanced interplay among specialized cell types transitioning between resting and stimulated states. Despite its importance, the regulatory landscape of this system has not yet been fully characterized. To address this gap, we collected ATAC-seq and RNA-seq data under resting and stimulated conditions for 25 immune cell types from peripheral blood of four healthy individuals, and seven cell types from three fetal thymus samples. We found that stimulation caused widespread chromatin remodeling, including a large class of response elements shared between stimulated B and T cells. Furthermore, several autoimmune traits showed significant heritability in stimulation-responsive elements from distinct cell types, highlighting the critical importance of these cell states in autoimmunity. Use of allele-specific read-mapping identified thousands of variants that alter chromatin accessibility in particular conditions. Notably, variants associated with changes in stimulation-specific chromatin accessibility were not enriched for associations with gene expression regulation in whole blood - a tissue commonly used in eQTL studies. Thus, large-scale maps of variants associated with gene regulation lack a condition important for understanding autoimmunity. As a proof-of-principle we identified variant rs6927172, which links stimulated T cell-specific chromatin dysregulation in the TNFAIP3 locus to ulcerative colitis and rheumatoid arthritis. Overall, our results provide a broad resource of chromatin landscape dynamics and highlight the need for large-scale characterization of effects of genetic variation in stimulated cells.

genomics

A quantitative and predictive model for RNA binding by human Pumilio proteins

High-throughput methodologies have enabled routine generation of RNA target sets and sequence motifs for RNA-binding proteins (RBPs). Nevertheless, quantitative approaches are needed to capture the landscape of RNA/RBP interactions responsible for cellular regulation. We have used the RNA-MaP platform to directly measure equilibrium binding for thousands of designed RNAs and to construct a predictive model for RNA recognition by the human Pumilio proteins PUM1 and PUM2. Despite prior findings of linear sequence motifs, our measurements revealed widespread residue flipping and instances of positional coupling. Application of our thermodynamic model to published in vivo crosslinking data reveals quantitative agreement between predicted affinities and in vivo occupancies. Our analyses suggest a thermodynamically driven, continuous Pumilio binding landscape that is negligibly affected by RNA structure or kinetic factors, such as displacement by ribosomes. This work provides a quantitative foundation for dissecting the cellular behavior of RBPs and cellular features that impact their occupancies.

biochemistry

High-resolution mapping of cancer cell networks using co-functional interactions

Powerful new technologies for perturbing genetic elements have expanded the study of genetic interactions in model systems ranging from yeast to human cell lines. However, technical artifacts can confound signal across genetic screens and limit the immense potential of parallel screening approaches. To address this problem, we devised a novel PCA-based method for eliminating these artifacts and bolstering sensitivity and specificity for detection of genetic interactions. Applying this strategy to a set of >300 whole genome CRISPR screens, we report ~1 million pairs of correlated \"co-functional\" genes that provide finer-scale information about cell compartments, biological pathways, and protein complexes than traditional gene sets. Lastly, we employed a gene community detection approach to implicate core genes for cancer growth and compress signal from functionally related genes in the same community into a single score. This work establishes new algorithms for probing cancer cell networks and motivates the acquisition of further CRISPR screen data across diverse genotypes and cell types to further resolve the complexity of cell signaling processes.

genomics

Large-scale, quantitative protein assays on a high-throughput DNA sequencing chip

High-throughput DNA sequencing techniques have enabled diverse approaches for linking DNA sequence to biochemical function. In contrast, assays of protein function have substantial limitations in terms of throughput, automation, and widespread availability. We have adapted an Illumina high-throughput sequencing chip to display an immense diversity of ribosomally-translated proteins and peptides, and then carried out fluorescence-based functional assays directly on this flow cell, demonstrating that a single, widely-available high-throughput platform can perform both sequencing-by-synthesis and protein assays. We quantified the binding of the M2 anti-FLAG antibody to a library of 1.3x104 variant FLAG peptides, exploring non-additive effects of combinations of mutations and discovering a \"superFLAG\" epitope variant. We also measured the enzymatic activity of 1.56x105 molecular variants of full-length of human O6-alkylguanine-DNA alkyltransferase (SNAP-tag). This comprehensive corpus of catalytic rates linked to amino acid sequence perturbations revealed amino acid interaction networks and cooperativity, linked positive cooperativity to structural proximity, and revealed ubiquitous positively-cooperative interactions with histidine residues.

molecular biology

RNA tertiary structure energetics predicted by an ensemble model of the RNA double helix

Over 50% of residues within functional structured RNAs are base-paired in Watson-Crick helices, but it is not fully understood how these helices geometric preferences and flexibility might influence RNA tertiary structure. Here, we show experimentally and computationally that the ensemble fluctuations of RNA helices substantially impact RNA tertiary structure stability. We updated a model for the conformational ensemble of the RNA helix using crystallographic structures of Watson-Crick base pair steps. To test this model, we made blind predictions of the thermodynamic stability of >1500 tertiary assemblies with differing helical sequences and compared calculations to independent measurements from a high-throughput experimental platform. The blind predictions accounted for thermodynamic effects from changing helix sequence and length with unexpectedly tight accuracies (RMSD of 0.34 and 0.77 kcal/mol, respectively). These comparisons lead to a detailed picture of how RNA base pair steps fluctuate within complex assemblies and suggest a new route toward predicting RNA tertiary structure formation and energetics.

biophysics

Nonparametric analysis of contributions to variance in genomics and epigenomics data

Functional genomics studies, despite increasingly varied assay types and complex experimental designs, are typically analyzed by methods that are unable to identify confounding effects and that incorporate parametric assumptions particular to gene expression data. We present MAVRIC, a nonparametric method to quantify variance explained by experimental covariates and perform differential analysis on arbitrary data types. We demonstrate that MAVRIC can accurately associate covariates with underlying data variance, deliver sensitive and specific identification of genomic loci with differential counts, and provide effective noise reduction of large-scale consortium data sets.

bioinformatics

Joint single-cell DNA accessibility and protein epitope profiling reveals environmental regulation of epigenomic heterogeneity

Here we introduce Protein-indexed Assay of Transposase Accessible Chromatin with sequencing (Pi-ATAC) that combines single-cell chromatin and proteomic profiling. In conjunction with DNA transposition, the levels of multiple cell surface or intracellular protein epitopes are recorded by index flow cytometry and positions in arrayed microwells, and then subject to molecular barcoding for subsequent pooled analysis. Pi-ATAC simultaneously identifies the epigenomic and proteomic heterogeneity in individual cells. Pi-ATAC reveals a casual link between transcription factor abundance and DNA motif access, and deconvolute cell types and states in the tumor microenvironment in vivo. We identify a dominant role for hypoxia, marked by HIF1 protein, in the tumor microvenvironment for shaping the regulome in a subset of epithelial tumor cells.

genomics

Neutralizing Gatad2a-Chd4-Mbd3 Axis within the NuRD Complex Facilitates Deterministic Induction of Naive Pluripotency

The Nucleosome Remodeling and Deacytelase (NuRD) complex is a co-repressive complex involved in many pathological and physiological processes in the cell. Previous studies have identified one of its components, Mbd3, as a potent inhibitor for reprogramming of somatic cells to pluripotency. Following OSKM induction, early and partial depletion of Mbd3 protein followed by applying naive ground-state pluripotency conditions, results in a highly efficient and near-deterministic generation of mouse iPS cells. Increasing evidence indicates that the NuRD complex assumes multiple mutually exclusive protein complexes, and it remains unclear whether the deterministic iPSC phenotype is the result of a specific NuRD sub complex. Since complete ablation of Mbd3 blocks somatic cell proliferation, here we aimed to identify alternative ways to block Mbd3-dependent NuRD activity by identifying additional functionally relevant components of the Mbd3/NuRD complex during early stages of reprogramming. We identified Gatad2a (also known as P66), a relatively uncharacterized NuRD-specific subunit, whose complete deletion does not impact somatic cell proliferation, yet specifically disrupts Mbd3/NuRD repressive activity on the pluripotency circuit during both stem cell differentiation and reprogramming to pluripotency. Complete ablation of Gatad2a in somatic cells, but not Gatad2b, results in a deterministic naive iPSC reprogramming where up to 100% of donor somatic cells successfully complete the process within 8 days. Genetic and biochemical analysis established a distinct sub-complex within the NuRD complex (Gatad2a-Chd4-Mbd3) as the functional and biochemical axis blocking reestablishment of murine naive pluripotency. Disassembly of this axis by depletion of Gatad2a, results in resistance to conditions promoting exit of naive pluripotency and delays differentiation. We further highlight context- and posttranslational dependent modifications of the NuRD complex affecting its interactions and assembly in different cell states. Collectively, our work unveils the distinct functionality, composition and interactions of Gatad2a-Chd4-Mbd3/NuRD subcomplex during the resolution and establishment of mouse naive pluripotency.

developmental biology

Prospects for recurrent neural network models to learn RNA biophysics from high-throughput data

RNA is a functionally versatile molecule that plays key roles in genetic regulation and in emerging technologies to control biological processes. Computational models of RNA secondary structure are well-developed but often fall short in making quantitative predictions of the behavior of multi-RNA complexes. Recently, large datasets characterizing hundreds of thousands of individual RNA complexes have emerged as rich sources of information about RNA energetics. Meanwhile, advances in machine learning have enabled the training of complex neural networks from large datasets. Here, we assess whether a recurrent neural network model, Ribonet, can learn from high-throughput binding data, using simulation and experimental studies to test model accuracy but also determine if they learned meaningful information about the biophysics of RNA folding. We began by evaluating the model on energetic values predicted by the Turner model to assess whether the neural network could learn a representation that recovered known biophysical principles. First, we trained Ribonet to predict the simulated free energy of an RNA in complex with multiple input RNAs. Our model accurately predicts free energies of new sequences but also shows evidence of having learned base pairing information, as assessed by in silico double mutant analysis. Next, we extended this model to predict the simulated affinity between an arbitrary RNA sequence and a reporter RNA. While these more indirect measurements precluded the learning of basic principles of RNA biophysics, the resulting model achieved sub-kcal/mol accuracy and enabled design of simple RNA input responsive riboswitches with high activation ratios predicted by the Turner model from which the training data were generated. Finally, we compiled and trained on an experimental dataset comprising over 600,000 experimental affinity measurements published on the Eterna open laboratory. Though our tests revealed that the model likely did not learn a physically realistic representation of RNA interactions, it nevertheless achieved good performance of 0.76 kcal/mol on test sets with the application of transfer learning and novel sequence-specific data augmentation strategies. These results suggest that recurrent neural network architectures, despite being naive to the physics of RNA folding, have the potential to capture complex biophysical information. However, more diverse datasets, ideally involving more direct free energy measurements, may be necessary to train de novo predictive models that are consistent with the fundamentals of RNA biophysics.\n\nAuthor SummaryThe precise design of RNA interactions is essential to gaining greater control over RNA-based biotechnology tools, including designer riboswitches and CRISPR-Cas9 gene editing. However, the classic model for energetics governing these interactions fails to quantitatively predict the behavior of RNA molecules. We developed a recurrent neural network model, Ribonet, to quantitatively predict these values from sequence alone. Using simulated data, we show that this model is able to learn simple base pairing rules, despite having no a priori knowledge about RNA folding encoded in the network architecture. This model also enables design of new switching RNAs that are predicted to be effective by the \"ground truth\" simulated model. We applied transfer learning to retrain Ribonet using hundreds of thousands of RNA-RNA affinity measurements and demonstrate simple data augmentation techniques that improve model performance. At the same time, data diversity currently available set limits on Ribonets accuracy. Recurrent neural networks are a promising tool for modeling nucleic acid biophysics and may enable design of complex RNAs for novel applications.

biophysics

High-Resolution Dissection of Conducive Reprogramming Trajectory to Ground State Pluripotency

The ability to reprogram somatic cells into induced pluripotent stem cells (iPSCs) with four transcription factors Oct4, Sox2, Klf4 and cMyc (abbreviated as OSKM)1 has provoked interest to define the molecular characteristics of this process2-7. Despite important progress, the dynamics of epigenetic reprogramming at high resolution in correctly reprogrammed iPSCs and throughout the entire process remain largely undefined. This gap in understanding results from the inefficiency of conventional reprogramming methods coupled with the difficulty of prospectively isolating the rare cells that eventually correctly reprogram into iPSCs. Here we characterize cell fate conversion from fibroblast to iPSC using a highly efficient deterministic murine reprogramming system engineered through optimized inhibition of Gatad2a-Mbd3/NuRD repressive sub-complex. This comprehensive characterization provides single-day resolution of dynamic changes in levels of gene expression, chromatin modifications, TF binding, DNA accessibility and DNA methylation. The integrative analysis identified two transcriptional modules that dominate successful reprogramming. One consists of genes whose transcription is regulated by on/off epigenetic switching of modifications in their promoters (abbreviated as ESPGs), and the second consists of genes with promoters in a constitutively active chromatin state, but a dynamic expression pattern (abbreviated as CAPGs). ESPGs are mainly regulated by OSK, rather than Myc, and are enriched for cell fate determinants and pluripotency factors. CAPGs are predominantly regulated by Myc, and are enriched for cell biosynthetic regulatory functions. We used the ESPG module to study the identity and temporal occurrence of activating and repressing epigenetic switching during reprogramming. Removal of repressive chromatin modifications precedes chromatin opening and binding of RNA polymerase II at enhancers and promoters, and the opposite dynamics occur during repression of enhancers and promoters. Genome wide DNA methylation analysis demonstrated that de novo DNA methylation is not required for highly efficient conducive iPSC reprogramming, and identified a group of super-enhancers targeted by OSK, whose early demethylation marks commitment to a successful reprogramming trajectory also in inefficient conventional reprogramming systems. CAPGs are distinctively regulated by multiple synergystic ways: 1) Myc activity, delivered either endogenously or exogenously, dominates CAPG expression changes and is indispensable for induction of pluripotency in somatic cells; 2) A change in tRNA codon usage which is specific to CAPGs, but not ESPGs, and favors their translation. In summary, our unbiased high-resolution mapping of epigenetic changes on somatic cells that are committed to undergo successful reprogramming reveals interleaved epigenetic and biosynthetic reconfigurations that rapidly commission and propel conducive reprogramming toward naive pluripotency.

developmental biology

Enhancer connectome in primary human cells reveals target genes of disease-associated DNA elements

The challenge of linking intergenic mutations to target genes has limited molecular understanding of diverse human diseases. Here, we show H3K27ac HiChIP generates high-resolution contact maps of active enhancers and target genes in rare primary human T cell subtypes and coronary artery smooth muscle cells. Differentiation of naive T cells to either T helper 17 cells or regulatory T cells create subtype-specific enhancer-promoter interactions, specifically at regions of shared DNA accessibility. These data provide a principled means of assigning molecular functions to autoimmune and cardiovascular disease risk variants, linking hundreds of noncoding variants to putative gene targets. Target genes identified with HiChIP are further supported by CRISPR interference and activation at linked enhancers, by the presence of expression quantitative trait loci, and by allele-specific enhancer loops in patient-derived primary cells. The majority of disease-associated enhancers contact genes beyond the nearest gene in the linear genome, leading to a four-fold increase of potential target genes for autoimmune and cardiovascular diseases.

genomics

An improved ATAC-seq protocol reduces background and enables interrogation of frozen tissues

We present Omni-ATAC, an improved ATAC-seq protocol for chromatin accessibility profiling that works across multiple applications with substantial improvement of signal-to-background ratio and information content. The Omni-ATAC protocol enables chromatin accessibility profiling from archival frozen tissue samples and 50 m sections, revealing the activities of disease-associated DNA elements in distinct human brain structures. The Omni-ATAC protocol enables the interrogation of personal regulomes in tissue context and translational studies.

genomics

Unsupervised Clustering And Epigenetic Classification Of Single Cells

Characterizing epigenetic heterogeneity at the cellular level is a critical problem in the modern genomics era. Assays such as single cell ATAC-seq (scATAC-seq) offer an opportunity to interrogate cellular level epigenetic heterogeneity through patterns of variability in open chromatin. However, these assays exhibit technical variability that complicates clear classification and cell type identification in heterogeneous populations. We present scABC, an R package for the unsupervised clustering of single cell epigenetic data, to classify scATAC-seq data and discover regions of open chromatin specific to cell identity.

bioinformatics

Chromatin-associated RNA sequencing (ChAR-seq) maps genome-wide RNA-to-DNA contacts

RNA is a critical component of chromatin in eukaryotes, both as a product of transcription, and as an essential constituent of ribonucleoprotein complexes that regulate both local and global chromatin states. Here we present a proximity ligation and sequencing method called Chromatin-Associated RNA sequencing (ChAR-seq) that maps all RNA-to-DNA contacts across the genome. ChAR-seq provides unbiased, de novo identification of targets of chromatin-bound RNAs including nascent transcripts, chromosome-specific dosage compensation ncRNAs, and genome-wide trans-associated RNAs involved in co-transcriptional RNA processing.

genomics

chromVAR: Inferring transcription factor variation from single-cell epigenomic data

Single cell ATAC-seq (scATAC) yields sparse data that makes application of conventional computational approaches for data analysis challenging or impossible. We developed chromVAR, an R package for analyzing sparse chromatin accessibility data by estimating the gain or loss of accessibility within sets of peaks sharing the same motif or annotation while controlling for known technical biases. chromVAR enables accurate clustering of scATAC-seq profiles and enables characterization of known, or the de novo identification of novel, sequence motifs associated with variation in chromatin accessibility across single cells or other sparse epigenomic data sets.

bioinformatics

Chromatin accessibility dynamics reveal novel functional enhancers in C. elegans

Chromatin accessibility, a crucial component of genome regulation, has primarily been studied in homogeneous and simple systems, such as isolated cell populations or early-development models. Whether chromatin accessibility can be assessed in complex, dynamic systems in vivo with high sensitivity remains largely unexplored. In this study, we use ATAC-seq to identify chromatin accessibility changes in a whole animal, the model organism C. elegans, from embryogenesis to adulthood. Chromatin accessibility changes between developmental stages are highly reproducible, recapitulate histone modification changes, and reveal key regulatory aspects of the epigenomic landscape throughout organismal development. We find that over 5,000 distal non-coding regions exhibit dynamic changes in chromatin accessibility between developmental stages, and could thereby represent putative enhancers. When tested in vivo, several of these putative enhancers indeed drive novel cell-type-and temporal-specific patterns of expression. Finally, by integrating transcription factor binding motifs in a machine learning framework, we identify EOR-1 as a unique transcription factor that may regulate chromatin dynamics during development. Our study provides a unique resource for C. elegans, a system in which the prevalence and importance of enhancers remains poorly characterized, and demonstrates the power of using whole organism chromatin accessibility to identify novel regulatory regions in complex systems.

genomics