bioRxiv ScienceSearch

Biology subjects

Ni, P.

Publications and source records attributed to Ni, P..

3 recordsLinked to original sources

Deciphering epigenomic code for cell differentiation using deep learning

Epigenomic markers, such as histone modifications, play important roles in cell fate determination and type maintenance during cell differentiation. Although genomic sequence plays a crucial role in establishing the unique epigenome in each cell type produced during cell differentiation, little is known about the sequence determinants that lead to the unique epigenomes of the cells. Here, using a dataset of six histone markers measured in four human CD4+ T cell types produced at different stages of T cell development, we showed that two types of highly accurate deep convolutional neural networks (CNNs) constructed for each cell type and for each histone marker are a powerful strategy to uncover the sequence determinants of the various histone modification patterns in difference cell types. We found that sequence motifs learned by the CNN models are highly similar to known binding motifs of transcription factors known to play important roles in CD4+ T cell differentiation. Our results suggest that both the unique histone modification patterns in each cell type and the different patterns of the same histone marker in different cell types are determined by a set of motifs with unique combinations. Interestingly, the level of shared few motifs learned in the different cell models reflect the lineage relationships of the cells, while the level of few shared motifs learned in different histone marker models reflect their functional relationships. Furthermore, using these models, we can predict the importance of the learned motifs and their interactions in determining specific histone marker patterns in the cell types.

bioinformatics

Ultra-fast and accurate motif finding in large ChIP-seq datasets reveals transcription factor binding patterns

The availability of a large volume of chromatin immunoprecipitation followed by sequencing (ChIP-seq) datasets for various transcription factors (TF) has provided an unprecedented opportunity to identify all functional TF binding motifs clustered in the enhancers in genomes. However, the progress has been largely hindered by the lack of a highly efficient and accurate tool that is fast enough to find not only the target motifs, but also cooperative motifs contained in very large ChIP-seq datasets with a binding peak length of typical enhancers ([~] 1,000 bp). To circumvent this hurdle, we herein present an ultra-fast and highly accurate motif-finding algorithm, ProSampler, with automatic motif length detection. ProSampler first identifies significant k-mers in the dataset and combines highly similar significant k-mers to form preliminary motifs. ProSampler then merges preliminary motifs with subtle similarity using a novel graph-based Gibbs sampler to find core motifs. Finally, ProSampler extends the core motifs by applying a two-proportion z-test to the flanking positions to identify motifs longer than k. As the number of preliminary motifs is much smaller than that of k-mers in a dataset, we greatly reduce the search space of the Gibbs sampler compared with conventional ones. By storing flanking sequences in a hash table, we avoid extensive IO and the necessity of examining all lengths of motifs in an interval. When evaluated on both synthetic and real ChIP-seq datasets, ProSampler runs orders of magnitude faster than the fastest existing tools while more accurately discovering primary motifs as well as cooperative motifs than do the best existing tools. Using ProSampler, we revealed previously unknown complex motif occurrence patterns in large ChIP-seq datasets, thereby providing insights into the mechanisms of cooperative TF binding for gene transcriptional regulation. Therefore, by allowing fast and accurate mining of the entire ChIP-seq datasets, ProSampler can greatly facilitate the efforts to identify the entire cis-regulatory code in genomes.

bioinformatics

DeepSignal: detecting DNA methylation state from Nanopore sequencing reads using deep-learning

The Oxford Nanopore sequencing enables to directly detect methylation sites in DNA from reads without extra laboratory techniques. In this study, we develop DeepSignal, a deep learning method to detect DNA methylated sites from Nanopore sequencing reads. DeepSignal construct features from both raw electrical signals and signal sequences in Nanopore reads. Testing on Nanopore reads of pUC19, E. coli and human, we show that DeepSignal can achieve both higher read level and genome level accuracy on detecting 6mA and 5mC methylation comparing to previous HMM based methods. Moreover, DeepSignal achieves similar performance cross different methylation bases and different methylation motifs. Furthermore, DeepSignal can detect 5mC and 6mA methylation states of genome sites with above 90% genome level accuracy under just 5X coverage using controlled methylation data.

bioinformatics