bioRxiv Science⌕ Search

Biology subjects

Kamulegeya, F.

Publications and source records attributed to Kamulegeya, F..

3 recordsLinked to original sources

Temporal regulation of a spatial patterning factor in Drosophila neurogenesis

A central question in neurobiology is how the transient programs that pattern neural progenitors are translated into the enormous, stable diversity of neuronal types. Spatial and temporal cues act only briefly, yet each neuron's identity is defined and maintained for life by terminal selector transcription factors (TFs). How a neuron's developmental origin is read out into a particular selector code remains poorly understood. Some current models propose that spatial and temporal origins are inherited independently through separate selectors. We show instead that, in the Drosophila optic lobe, the same selector can be activated by different patterning axes through physically distinct enhancers, even within the same lineage. Visual system homeobox (Vsx1) spatially patterns a central neuroepithelial domain and later acts as a terminal selector in dozens of neuronal types, most originating exclusively from that domain. However, in Dm2 neurons that are produced from every domain, it is regulated not by neuroepithelial Vsx1 but by the neuroblast temporal TF BarH1, through an enhancer distinct from its domain-specific ones. Combining in vivo reporters with sequence-to-accessibility deep-learning models, we identify and disrupt the key binding sites in this enhancer, impairing its Dm2-specific activity. Reciprocally, the temporal TF Homeobrain (Hbn) acts as a terminal selector in the related neuron Mi21 independently of its neuroblast temporal window: its expression in these late-born neurons is instead placed under dorsoventral spatial control. Patterning inputs therefore need not be partitioned across separate selectors but converge combinatorially on the modular enhancers of shared ones, revealing a cis-regulatory logic that re-encodes this limited set of inputs into vast neuronal diversity.

neuroscience↗

PISA: a versatile interpretation tool for visualizing cis-regulatory rules in genomic data

Sequence-to-function neural networks learn cis-regulatory sequence rules driving many types of genomic data. Interpreting these models to relate the sequence rules to underlying biological processes remains challenging, especially for complex genomic readouts such as MNase-seq, which maps nucleosome occupancy but is confounded by experimental bias. We introduce pairwise influence by sequence attribution (PISA), an interpretation tool that combinatorially decodes which bases contributed to the readout at a specific genomic coordinate. PISA visualizes the effects of transcription factor motifs, detects undiscovered motifs with complex contribution patterns, and reveals experimental biases. By learning the bias for MNase-seq, PISA enables unprecedented nucleosome prediction models, allowing the de novo discovery of nucleosome-positioning motifs and their longrange chromatin effects, as well as the design of sequences with altered nucleosome configurations. These results show that PISA is a versatile tool that expands our ability to train and interpret sequence-to-function neural networks on genomics data and understand the underlying cis-regulatory code.

genomics↗

Interpreting the CTCF-mediated sequence grammar of genome folding with AkitaV2

Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful, (ii) insulation strength is highly dependent on sequences flanking the core CTCF binding motif, (iii) core and flanking sequences are broadly compatible, and (iv) core and flanking nucleotides contribute largely additively to overall strength. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries. Author SummaryMammalian genomes are spatially organized in 3D with profound consequences for all processes involving DNA. CTCF is a key genome organizer, recognizing numerous sites and creating a variety of contact patterns across the genome. Despite the importance of CTCF, the sequence determinants and grammar of how individual sites collectively instruct genome folding remain unclear. This work leverages the ability of Akita, a deep neural network, to make high-throughput predictions for genome folding after DNA sequence perturbations. Using Akita, we make several experimentally testable predictions. First, only a minority of annotated sites individually impact folding, and flanking DNA sequences greatly modulate their impact. Second, multiple sites together influence folding based on their number, orientation, and spacing. In sum, we provide a roadmap for interpreting neural networks to better understand genome folding and important considerations for the design of experiments.

genomics↗