bioRxiv Science⌕ Search

Biology subjects

Nobusada, T.

Publications and source records attributed to Nobusada, T..

3 recordsLinked to original sources

The human RNA-DNA interactome is cell type-specific and dynamic

More than twenty years ago, the FANTOM consortium uncovered that mammalian genomes are pervasively transcribed, revealing multitudes of RNAs with unknown functions. A subset of these transcripts has since then been linked to transcriptional control and to chromatin organization via their ability to interact with DNA, suggesting that chromatin-associated RNAs could be key players in genome regulation. Although recent technological advances now enable the mapping of genome-wide RNA-DNA contacts, a lack of analyses integrating these methods with other genomic features and across multiple cellular contexts hinders our comprehensive understanding of the principles underlying RNA-DNA interactions and of their biological importance. As part of the FANTOM6 project, we thus generated RNA-DNA interaction maps in 16 different human cell types, then combined these contacts with multiple layers of other genomic data to investigate how patterns of interaction between RNA and DNA relate to chromatin organization and function. We show that the RNA-DNA interactome is highly dynamic yet reproducibly organized in cell-type specific networks, constituted of a great diversity of interactions that vary in function of their distance, the nature of their sources and the chromatin state of their targets. In particular, we detected numerous regulatory elements that exhibit marked changes in activity when differentially bound by transcripts, implying that thousands of RNA-DNA interactions can play a mechanistic role in gene expression. This regulatory function correlates with RNA-protein interactions and significantly associates with cell type-relevant and disorder-related traits. In addition to providing essential resources for future research in RNA-mediated chromatin regulation, cellular biology and human diseases, our study thus establishes the RNA-DNA interactome as a new genome regulatory layer that defines and maintains cellular identity and behavior.

genomics↗

A genome-wide, machine learning-guided exploration of the cis-regulatory code involved in neuronal differentiation

Gene expression is controlled by proximal and distal cis-regulatory elements (CREs), containing DNA motifs bound by various transcription factors (TFs). Other sequence features, such as specific k-mers or low complexity regions, have also been implicated [1-3]. However, in a dynamic biological process such as cell differentiation, we lack an understanding of how the transcriptional activity of CREs progressively change and what sequence features underlie these transitions, which may reflect common and/or coordinated regulatory processes. Here, we use single-cell ATAC-seq and RNA-seq to follow, at a genome scale, CREs along differentiation of induced pluripotent stem cells into cortical neurons and develop a method to automatically identify the diversity of CRE profiles and their underlying sequence features. We propose a machine-learning guided clustering algorithm, STOIC (Statistical learning TO Inform Clustering), that jointly learns an unsupervised clustering of the CREs in the space of the activity profiles and a supervised predictor associated with each cluster in the DNA-sequence space. STOIC is specifically designed to provide readily interpretable results. We show that the method identifies CRE profiles associated with highly predictive sequence features and outperforms methods solely concerned with co-activity clustering on this task. Orthogonal data collected in the same settings link the inferred CRE clusters to specific enhancer or promoter signatures. Furthermore, we show that the DNA features unveiled by STOIC reflect biologically relevant regulators and offer a valuable basis to dissect elements of the cisregulatory grammar. Finally, we demonstrate the general applicability of STOIC by analyzing five bulk CAGE datasets of human cells responding to various treatments.

bioinformatics↗

CFC-seq: identification of full-length capped RNAs unveil enhancer-derived transcription

Long-read sequencing has transformed transcriptome profiling, yet capturing full-length, non-polyadenylated transcripts like enhancer RNAs (eRNAs) remains challenging. Here, we introduce CFC-seq, combining cap-trapping and in vitro poly(A)-tailing to sequence poly(A) and non-poly(A) RNAs with precise transcription start site. Paired with our assembler, SALA, we identified 39,425 novel transcriptional units, including [~]24,000 eRNAs. Our data reveal a distinct genomic code governing eRNA fate dictated by core promoter architecture. CpG-island enhancers show high chromatin connectivity but yield short, exosome-sensitive RNAs. Conversely, TATA-box enhancers systematically co-opt LTR retrotransposons to inherit structural motifs that produce long, stable, and spliced RNAs. Mechanistically, the pioneer factor NF-Y activates these viral elements to license transcription, balanced by TEAD4 activity across a dual-gear regulatory axis. Finally, non-poly(A) eRNAs terminate via exosome-associated processing at structural-depleted cleavage zones. This comprehensive annotation links enhancer sequence architecture to RNA fate, providing a new transformative framework for decoding the functional human genome. HighlightsO_LIExpanded genomic architecture: CFC-seq unmasks a hidden layer of human transcriptome, identifying 39,425 novel transcriptional units with high-confidence TSS support, including [~]24,000 eRNAs. C_LIO_LITSS-first assembler: We introduce SALA, a specialized long-read assembler that prioritizes authentic 5 Cap-trapped ends to accurately reconstruct the TSS-resolved transcript models. C_LIO_LIGenomic code of eRNA fate: CGI enhancers drive short and exosome-sensitive transcripts associated with repressive H3K27me3 mark and high chromatin connectivity. TATA-box enhancers produce cell-type-specific, long, stable, and frequently spliced eRNAs. C_LIO_LIEvolutionary co-option of retrotransposons: A major fraction of TATA-box eRNAs originate from LTR retrotransposons, providing a direct mechanism for integration of viral elements into the human regulatory landscape. C_LIO_LIA dual-gear pioneering axis: The pioneer factor NF-Y activates unprimed LTR-TATA enhancers to license transcription independent of histone acetylation cascades, operating in parallel with TEAD4-mediated activation. C_LIO_LIStructural determinants of eRNA termination: Non-poly(A) eRNA TES features a secondary structure depletion zone that coordinates pol II termination and calibrates exosome-mediated turnover. C_LI

genomics↗