bioRxiv ScienceSearch

Biology subjects

Kaplan, T.

Publications and source records attributed to Kaplan, T..

5 recordsLinked to original sources

Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease

Methylation patterns of circulating cell-free DNA (cfDNA) contain rich information about recent cell death events in the body. Here, we present an approach for unbiased determination of the tissue origins of cfDNA, using a reference methylation atlas of 25 human tissues and cell types. The method is validated using in silico simulations as well as in vitro mixes of DNA from different tissue sources at known proportions. We show that plasma cfDNA of healthy donors originates from white blood cells (55%), erythrocyte progenitors (30%), vascular endothelial cells (10%) and hepatocytes (1%). Deconvolution of cfDNA from patients reveals tissue contributions that agree with clinical findings in sepsis, islet transplantation, cancer of the colon, lung, breast and prostate, and cancer of unknown primary. We propose a procedure which can be easily adapted to study the cellular contributors to cfDNA in many settings, opening a broad window into healthy and pathologic human tissue dynamics.

molecular biology

Enhancer Identification using Transfer and Adversarial Deep Learning of DNA Sequences

Enhancer sequences regulate the expression of genes from afar by providing a binding platform for transcription factors, often in a tissue-specific or context-specific manner. Despite their importance in health and disease, our understanding of these DNA sequences, and their regulatory grammar, is limited. This impairs our ability to identify new enhancers along the genome, or to understand the effect of enhancer mutations and their role in genetic diseases.\n\nWe trained deep Convolutional Neural Networks (CNN) to identify enhancer sequences in multiple species. We used multiple biological datasets, including simulated sequences, in vivo binding data of single transcription factors and genome-wide chromatin maps of active enhancers in 17 mammalian species. Our deep networks obtained high classification accuracy by combining two training strategies: First, training on enhancers vs. non-enhancer background sequences, we identified short (1-4bp) low-complexity motifs. Second, by replacing the negative training set by adversarial k-order random shuffles of enhancer sequences (thus maintaining base composition while shuttering longer motifs, including transcription factor binding sites), we identified a set of biologically meaningful motifs, unique to enhancers. In addition, classification performance improved when combining positive data from all species together, showing a shared mammalian regulatory architecture.\n\nOur results demonstrate that design of adversarial training data, and transfer of learned parameters between networks trained on different species/datasets improve the overall performance and capture biologically meaningful information in the parameters of the learned network.\n\nContact: or.zuk@mail.huji.ac.il, tommy@cs.huji.ac.il

bioinformatics

GEM: A manifold learning based framework for reconstructing spatial organizations of chromosomes

Decoding the spatial organizations of chromosomes has crucial implications for studying eukaryotic gene regulation. Recently, Chromosomal conformation capture based technologies, such as Hi-C, have been widely used to uncover the interaction frequencies of genomic loci in high-throughput and genome-wide manner and provide new insights into the folding of three-dimensional (3D) genome structure. In this paper, we develop a novel manifold learning framework, called GEM (Genomic organization reconstructor based on conformational Energy and Manifold learning), to elucidate the underlying 3D spatial organizations of chromosomes from Hi-C data. Unlike previous chromatin structure reconstruction methods, which explicitly assume specific relationships between Hi-C interaction frequencies and spatial distances between distal genomic loci, GEM is able to reconstruct an ensemble of chromatin conformations by directly embedding the neigh-boring affinities from Hi-C space into 3D Euclidean space based on a manifold learning strategy that considers both the fitness of Hi-C data and the biophysical feasibility of the modeled structures, which are measured by the conformational energy derived from our current biophysical knowledge about the 3D polymer model. Extensive validation tests on both simulated interaction frequency data and experimental Hi-C data of yeast and human demonstrated that GEM not only greatly outperformed other state-of-art modeling methods but also reconstructed accurate chromatin structures that agreed well with the hold-out or independent Hi-C data and sparse geometric restraints derived from the previous fluorescence in situ hybridization (FISH) studies. In addition, as GEM can generate accurate spatial organizations of chromosomes by integrating both experimentally-derived spatial contacts and conformational energy, we for the first time extended our modeling method to recover long-range genomic interactions that are missing from the original Hi-C data. All these results indicated that GEM can provide a physically and physiologically valid 3D representations of the organizations of chromosomes and thus serve as an effective and useful genome structure reconstructor.

bioinformatics

Genome-wide Search for Zelda-like Chromatin Signatures Identifies GAF as a Pioneer Factor in Early Fly Development

MotivationThe protein Zelda was shown to play a key role in early Drosophila development, binding thousands of promoters and enhancers prior to maternal-to-zygotic transition (MZT), and marking them for transcriptional activation. Recently, we showed that Zelda acts through specific chromatin patterns of histone modifications to mark developmental enhancers and active promoters. Intriguingly, some Zelda sites still maintain these chromatin patterns in Drosophila embryos lacking maternal Zelda protein. This suggests that additional Zelda-like pioneer factors may act in early fly embryos.\n\nResultsWe developed a computational method to analyze and refine the chromatin landscape surrounding early Zelda peaks, using a multi-channel spectral clustering. This allowed us to characterize their chromatin patterns through MZT (mitotic cycles 8-14). Specifically, we focused on H3K4me1, H3K4me3, H3K18ac, H3K27ac, and H3K27me3 and identified three different classes of chromatin signatures, matching \"promoters\", \"enhancers\" and \"transiently bound\" Zelda peaks.\n\nWe then further scanned the genome using these chromatin patterns and identified additional loci - with no Zelda binding - that show similar chromatin patterns, resulting with hundreds of Zelda-independent putative enhancers. These regions were found to be enriched with GAGA factor (GAF, Trl), and are typically located near early developmental zygotic genes. Overall our analysis suggests that GAF, together with Zelda, plays an important role in activating the zygotic genome.\n\nAs we show, our computational approach offers an efficient algorithm for characterizing chromatin signatures around some loci of interest, and allows a genome-wide identification of additional loci with similar chromatin patterns.\n\nContact: tommy@cs.huji.ac.il

bioinformatics

Promoter-Enhancer Interactions Identified from Hi-C Data using Probabilistic Models and Hierarchical Topological Domains

Proximity-ligation methods as Hi-C allow us to map physical DNA-DNA interactions along the genome, and reveal its organization in topologically associating domains (TADs). As Hi-C data accumulate, computational methods were developed for identifying domain borders in multiple cell types and organisms.\n\nHere, we present PSYCHIC, a computational approach for analyzing Hi-C data and identifying Promoter-Enhancer interactions. We use a unified probabilistic model to segment the genome into domains, which we merge hierarchically and fit the Hi-C interaction map with a local background model. This allows us to estimate the expected number of interactions for every DNA-DNA pair, thus identifying over-represented interactions across the genome.\n\nBy analyzing published Hi-C data in human and mouse, we identified hundreds of thousands of putative enhancers and their target genes in multiple cell types, and compiled an extensive genome-wide catalog of gene regulation in human and mouse.

bioinformatics