bioRxiv Science⌕ Search

Biology subjects

Patel, Z. M.

Publications and source records attributed to Patel, Z. M..

8 recordsLinked to original sources

Cross-platform DNA motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors

A DNA sequence pattern, or "motif", is an essential representation of DNA-binding specificity of a transcription factor (TF). Any particular motif model has potential flaws due to shortcomings of the underlying experimental data and computational motif discovery algorithm. As a part of the Codebook/GRECO-BIT initiative, here we evaluated at large scale the cross-platform recognition performance of positional weight matrices (PWMs), which remain popular motif models in many practical applications. We applied ten different DNA motif discovery tools to generate PWMs from the "Codebook" data comprised of 4,237 experiments from five different platforms profiling the DNA-binding specificity of 394 human proteins, focusing on understudied transcription factors of different structural families. For many of the proteins, there was no prior knowledge of a genuine motif. By benchmarking-supported human curation, we constructed an approved subset of experiments comprising about 30% of all experiments and 50% of tested TFs which displayed consistent motifs across platforms and replicates. We present the Codebook Motif Explorer (https://mex.autosome.org), a detailed online catalog of DNA motifs, including the top-ranked PWMs, and the underlying source and benchmarking data. We demonstrate that in the case of high-quality experimental data, most of the popular motif discovery tools detect valid motifs and generate PWMs, which perform well both on genomic and synthetic data. Yet, for each of the algorithms, there were problematic combinations of proteins and platforms, and the basic motif properties such as nucleotide composition and information content offered little help in detecting such pitfalls. By combining multiple PMWs in decision trees, we demonstrate how our setup can be readily adapted to train and test binding specificity models more complex than PWMs. Overall, our study provides a rich motif catalog as a solid baseline for advanced models and highlights the power of the multi-platform multi-tool approach for reliable mapping of DNA binding specificities. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/619379v2_ufig1.gif" ALT="Figure 1"> View larger version (61K): org.highwire.dtl.DTLVardef@79561forg.highwire.dtl.DTLVardef@54c0aorg.highwire.dtl.DTLVardef@1c33f34org.highwire.dtl.DTLVardef@16a93ba_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical AbstractC_FLOATNO C_FIG

bioinformatics↗

Perspectives on Codebook: sequence specificity of uncharacterized human transcription factors

Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1, and even for well-characterized TFs, it remains controversial to what degree motifs accurately reflect binding sites in living cells2,3. Here, we describe a systematic effort ("Codebook") to determine the sequence specificity of 332 putative and poorly characterized human TFs. Over 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177, or 53%), of which most are unique to a single protein, thereby extending the vocabulary of sequence recognition encoded by human TFs by [~]100 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched within cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved, and direct TF binding sites across the human genome. These sites are concentrated in promoter regions, and are predictive of gene expression, illustrating that this new data atlas provides an important step forward in decoding the human genome.

genomics↗

Extensive binding of uncharacterized human transcription factors to genomic dark matter

The functional impact of a large portion of the human genome known as "dark matter DNA", which is composed mainly of repeat sequences, remains enigmatic. The genome also encodes hundreds of putative and poorly characterized transcription factors (TFs). Here, we determined genomic binding locations of 166 poorly characterized human TFs in living cells. Nearly half of them associate strongly with known regulatory regions such as promoters and enhancers, frequently co-localizing with each other at conserved motif matches. The other half often associate with genomic dark matter, however, at largely non-overlapping (i.e., unique) sites, via intrinsic sequence recognition. Fifty-four of the latter half, which we term "Dark TFs", mainly bind within regions of closed chromatin, with each recognizing a unique set of repeat sequences. The Dark TFs include many KZNFs, which are known to bind and silence TEs, and other TFs with apparent repressive functions. By contrast, some may be pioneers: we find that induction of TPRX1, a known regulator of zygotic preimplantation, leads to chromatin opening at many of its binding sites in the dark matter genome. Altogether, our results shed light on a large fraction of poorly characterized human TFs and simultaneously illuminate the diversity of function within the dark matter genome.

genomics↗

CRISPR-CLEAR: Nucleotide-Resolution Mapping of Regulatory Elements via Allelic Readout of Tiled Base Editing

CRISPR tiling screens have advanced the identification and characterization of regulatory sequences but are limited by low resolution arising from the indirect readout of editing via guide RNA sequencing. This study introduces CRISPR-CLEAR, an end-to-end experimental assay and computational pipeline, which leverages targeted sequencing of CRISPR-introduced alleles at the endogenous target locus following dense base-editing mutagenesis. This approach enables the dissection of regulatory elements at nucleotide resolution, facilitating a direct assessment of genotype-phenotype effects.

genomics↗

DNA-Diffusion: Leveraging Generative Models for Controlling Chromatin Accessibility and Gene Expression via Synthetic Regulatory Elements

The challenge of systematically modifying and optimizing regulatory elements for precise gene expression control is central to modern genomics and synthetic biology. Advancements in generative AI have paved the way for designing synthetic sequences with the aim of safely and accurately modulating gene expression. We leverage diffusion models to design context-specific DNA regulatory sequences, which hold significant potential toward enabling novel therapeutic applications requiring precise modulation of gene expression. Our framework uses a cell type-specific diffusion model to generate synthetic 200 bp regulatory elements based on chromatin accessibility across different cell types. We evaluate the generated sequences based on key metrics to ensure they retain properties of endogenous sequences: transcription factor binding site composition, potential for cell type-specific chromatin accessibility, and capacity for sequences generated by DNA diffusion to activate gene expression in different cell contexts using state-of-the-art prediction models. Our results demonstrate the ability to robustly generate DNA sequences with cell type-specific regulatory potential. DNA-Diffusion paves the way for revolutionizing a regulatory modulation approach to mammalian synthetic biology and precision gene therapy.

synthetic biology↗

Benchmarking computational methods to identify spatially variable genes and peaks

Spatially resolved transcriptomics offers unprecedented insight by enabling the profiling of gene expression within the intact spatial context of cells, effectively adding a new and essential dimension to data interpretation. To efficiently detect spatial structure of interest, an essential step in analyzing such data involves identifying spatially variable genes. Despite researchers having developed several computational methods to accomplish this task, the lack of a comprehensive benchmark evaluating their performance remains a considerable gap in the field. Here, we present a systematic evaluation of 14 methods using 60 simulated datasets generated by four different simulation strategies, 12 real-world transcriptomics, and three spatial ATAC-seq datasets. We find that spatialDE2 consistently outperforms the other benchmarked methods, and Morans I achieves competitive performance in different experimental settings. Moreover, our results reveal that more specialized algorithms are needed to identify spatially variable peaks.

bioinformatics↗

Immune-Epithelial Dynamics and Tissue Remodeling in Chronically Inflamed Nasal Epithelium via Multi-scaled Transcriptomics

Chronic rhinosinusitis (CRS) is a common inflammatory disease of the sinonasal cavity that affects millions of individuals worldwide. The complex pathophysiology of CRS remains poorly understood, with emerging evidence implicating the orchestration between diverse immune and epithelial cell types in disease progression. We applied single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics to both dissociated and intact, freshly isolated sinonasal human tissues to investigate the cellular and molecular heterogeneity of CRS with and without nasal polyp formation compared to non-CRS control samples. Our findings reveal a mechanism for macrophage-eosinophil recruitment into the nasal mucosa, systematic dysregulation of CD4+ and CD8+ T cells, and enrichment of mast cell populations to the upper airway tissues with intricate interactions between mast cells and CD4 T cells. Additionally, we identify immune-epithelial interactions and dysregulation, particularly involving understudied basal progenitor cells and Tuft chemosensory cells. We further describe a distinct basal cell differential trajectory in CRS patients with nasal polyps (NP), and link it to NP formation through immune-epithelial remodeling. By harnessing stringent patient tissue selection and advanced technologies, our study unveils novel aspects of CRS pathophysiology, and sheds light onto both intricate immune and epithelial cell interactions within the disrupted CRS tissue microenvironment and promising targets for therapeutic intervention. These findings expand upon existing knowledge of nasal inflammation and provide a comprehensive resource towards understanding the cellular and molecular mechanisms underlying this uniquely complex disease entity, and beyond.

immunology↗

Multi-species analysis of inflammatory response elements reveals ancient and lineage-specific contributions of transposable elements to NF-κB binding

Transposable elements (TEs) provide a source of transcription factor binding sites that can rewire conserved gene regulatory networks. NF-{kappa}B is an evolutionary conserved transcription factor complex primarily involved in innate immunity and inflammation. The extent to which TEs have contributed to NF-{kappa}B responses during mammalian evolution is not well established. Here we performed a multi-species analysis of TEs bound by the NF-{kappa}B subunit RELA (also known as p65) in response to the proinflammatory cytokine TNF. By comparing RELA ChIP-seq data from TNF-stimulated primary aortic endothelial cells isolated from human, mouse and cow, we found that 55 TE subfamilies were associated with RELA bound regions. These RELA-bound transposons possess active epigenetic features and reside near TNF-responsive genes. A prominent example of lineage-specific contribution of transposons comes from the bovine SINE subfamilies Bov-tA1/2/3 which collectively contributed over 14,000 RELA bound regions in cow. By comparing RELA binding data across species, we also found several examples of RELA motif-bearing TEs that colonized the genome prior to the divergence of the three species and contributed to species-specific RELA binding. For example, we found human RELA bound MER81 instances were enriched for the interferon gamma pathway and demonstrated that one RELA bound MER81 element can control the TNF-induced expression of Interferon Gamma Receptor 2 (IFNGR2). Using ancestral reconstructions, we found that RELA containing MER81 instances rapidly decayed during early primate evolution (> 50 million years ago (MYA)) before stabilizing since the separation of Old World monkeys (< 50 MYA). Taken together, our results suggest ancient and lineage-specific transposon subfamilies contributed to mammalian NF-{kappa}B regulatory networks.

genomics↗