bioRxiv Science⌕ Search

Biology subjects

Kempynck, N.

Publications and source records attributed to Kempynck, N..

5 recordsLinked to original sources

HyDrop v2: Scalable atlas construction for training sequence-to-function models

Deciphering cis-regulatory logic underlying cell type identity is a fundamental question in biology. Single-cell chromatin accessibility (scATAC-seq) data has enabled training of sequence-to-function deep learning models allowing decoding of enhancer logic and design of synthetic enhancers. Training such models requires large amounts of high-quality training data across species, organs, development, aging, and disease. To facilitate the cost-effective generation of large scATAC-seq atlases for model training, we developed a new version of the open-source microfluidic system HyDrop with increased sensitivity and scale: HyDrop v2. We generated HyDrop v2 atlases for the mouse cortex and Drosophila embryo development and compared them to atlases generated on commercial platforms. HyDrop v2 data integrates seamlessly with commercially available chromatin accessibility methods (10x Genomics). Differentially accessible regions and motif enrichment across cell types are equivalent between HyDrop-v2 and 10x atlases. Sequence-to-function models trained on either atlas are comparable as well in terms of enhancer predictions, sequence explainability, and transcription factor footprinting. By offering accessible data generation, enhancer models trained on HyDrop-v2 and mixed atlases can contribute to unraveling cell-type specific regulatory elements in health and disease.

bioinformatics↗

CREsted: modeling genomic and synthetic cell type-specific enhancers across tissues and species

Sequence-based deep learning models have become the state of the art for the analysis of the genomic regulatory code. Particularly for transcriptional enhancers, deep learning models excel at deciphering sequence features and grammar that underlie their spatiotemporal activity. To enable end-to-end enhancer modeling and design, we developed a software and modeling package, called CREsted. It combines preprocessing starting from single-cell ATAC-seq data; modeling with a choice of several architectures for training classification and regression models on either topics or pseudobulk peak heights; sequence design using multiple strategies; and downstream analysis through a collection of tools to locate transcription factor (TF) binding sites, infer the effect of a TF (activating or repressing) on enhancer accessibility, decipher enhancer grammar, and score gene loci. We demonstrate CREsted using a mouse cortex model that we validate using the BICCN collection of in vivo validated mouse brain enhancers. Classical enhancers in immune cells, including the IFNB1 enhanceosome are revisited using a PBMC model, and we assess the accuracy of TF binding site predictions with ChIP-seq. Additionally, we use CREsted to compare mesenchymal-like cancer cell states between tumor types; and we investigate different fine-tuning strategies of Borzoi within CREsted, comparing their performance and explainability with CREsted models trained from scratch. Finally, we train a CREsted model on a scATAC-seq atlas of zebrafish development and use this to design and in vivo validate cell type-specific synthetic enhancers in three tissues. For varying datasets, we demonstrate that CREsted facilitates efficient training and analyses, enabling scrutinization of the enhancer logic and design of synthetic enhancers across tissues and species. CREsted is available at https://crested.readthedocs.io.

genomics↗

The evolution of gene regulation in mammalian cerebellum development

Gene regulatory changes are considered major drivers of evolutionary innovations, including the cerebellums expansion during human evolution, yet they remain largely unexplored. In this study, we combined single-nucleus measurements of gene expression and chromatin accessibility from six mammals (human, bonobo, macaque, marmoset, mouse, and opossum) to uncover conserved and diverged regulatory networks in cerebellum development. We identified core regulators of cell identity and developed sequence-based models that revealed conserved regulatory codes. By predicting chromatin accessibility across 240 mammalian species, we reconstructed the evolutionary histories of human cis-regulatory elements, identifying sets associated with positive selection and gene expression changes, including the recent gain of THRB expression in cerebellar progenitor cells. Collectively, our work reveals the shared and mammalian lineage-specific regulatory programs governing cerebellum development.

evolutionary biology↗

Evaluating Methods for the Prediction of Cell Type-Specific Enhancers in the Mammalian Cortex

Identifying cell type-specific enhancers in the brain is critical to building genetic tools for investigating the mammalian brain. Computational methods for functional enhancer prediction have been proposed and validated in the fruit fly and not yet the mammalian brain. We organized the Brain Initiative Cell Census Network (BICCN) Challenge: Predicting Functional Cell Type-Specific Enhancers from Cross-Species Multi-Omics to assess machine learning and feature-based methods designed to nominate enhancer DNA sequences to target cell types in the mouse cortex. Methods were evaluated based on in vivo validation data from hundreds of cortical cell type-specific enhancers that were previously packaged into individual AAV vectors and retro-orbitally injected into mice. We find that open chromatin was a key predictor of functional enhancers, and sequence models improved prediction of non-functional enhancers that can be deprioritized as opposed to pursued for in vivo testing. Sequence models also identified cell type-specific transcription factor codes that can guide designs of in silico enhancers. This community challenge establishes a benchmark for enhancer prioritization algorithms and reveals computational approaches and molecular information that are crucial for identifying functional enhancers in mammalian cortical cell types. The results of this challenge bring us closer to understanding the complex gene regulatory landscape of the mammalian cortex and to designing more efficient genetic tools to target cortical cell types.

genomics↗

Enhancer-driven cell type comparison reveals similarities between the mammalian and bird pallium

Combinations of transcription factors govern the identity of cell types, which is reflected by enhancer codes in cis-regulatory genomic regions. Cell type-specific enhancer codes at nucleotide-level resolution have not yet been characterized for the mammalian neocortex. It is currently unknown whether these codes are conserved in other vertebrate brains, and whether they are informative to resolve homology relationships for species that lack a neocortex such as birds. To compare enhancer codes of cell types from the mammalian neocortex with those from the bird pallium, we generated single-cell multiome and spatially-resolved transcriptomics data of the chicken telencephalon. We then trained deep learning models to characterize cell type-specific enhancer codes for the human, mouse, and chicken telencephalon. We devised three metrics that exploit enhancer codes to compare cell types between species. Based on these metrics, non-neuronal and GABAergic cell types show a high degree of regulatory similarity across vertebrates. Proposed homologies between mammalian neocortical and avian pallial excitatory neurons are still debated. Our enhancer code based comparison shows that excitatory neurons of the mammalian neocortex and the avian pallium exhibit a higher degree of divergence than other cell types. In contrast to existing evolutionary models, the mammalian deep layer excitatory neurons are most similar to mesopallial neurons; and mammalian upper layer neurons to hyper- and nidopallial neurons based on their enhancer codes. In addition to characterizing the enhancer codes in the mammalian and avian telencephalon, and revealing unexpected correspondences between cell types of the mammalian neocortex and the chicken pallium, we present generally applicable deep learning approaches to characterize and compare cell types across species via the genomic regulatory code.

genomics↗