bioRxiv ScienceSearch

Biology subjects

Chen, H.

Publications and source records attributed to Chen, H..

50 records · Page 3Linked to original sources

Phytophthora methylomes modulated by expanded 6mA methyltransferases are associated with adaptive genome regions

Filamentous plant pathogen genomes often display a bipartite architecture with gene sparse, repeat-rich compartments serving as a cradle for adaptive evolution. However, the extent to which this \"two-speed\" genome architecture is associated with genome-wide epigenetic modifications is unknown. Here, we show that the oomycete plant pathogens Phytophthora infestans and Phytophthora sojae possess functional adenine N6- methylation (6mA) methyltransferases that modulate patterns of 6mA marks across the genome. In contrast, 5-methylcytosine (5mC) could not be detected in the two Phytophthora species. Methylated DNA IP Sequencing (MeDIP-seq) of each species revealed that 6mA is depleted around the transcriptional starting sites (TSS) and is associated with low expressed genes, particularly transposable elements. Remarkably, genes occupying the gene-sparse regions have higher levels of 6mA compared to the remainder of both genomes, possibly implicating the methylome in adaptive evolution of Phytophthora. Among three putative adenine methyltransferases, DAMT1 and DAMT3 displayed robust enzymatic activities. Surprisingly, single knockouts of each of the 6mA methyltransferases in P. sojae significantly reduced in vivo 6mA levels, indicating that the three enzymes are not fully redundant. MeDIP-seq of the damt3 mutant revealed uneven patterns of 6mA methylation across genes, suggesting that PsDAMT3 may have a preference for gene body methylation after the TSS. Our findings provide evidence that 6mA modification is an epigenetic mark of Phytophthora genomes and that complex patterns of 6mA methylation by the expanded 6mA methyltransferases may be associated with adaptive evolution in these important plant pathogens.

molecular biology

Parental allele-specific genome architecture and transcription during the cell cycle

A normal human somatic cell inherits two haploid genomes. Individual chromosomes of each pair have distinct parental origins and parental alleles are known to unequally contribute to cellular function. We integrated chromosome conformation (form) and gene transcription (function) analyses to dissect the dynamics of the maternal and paternal genomes in lymphoblastoid cells during the cell cycle. We found a distinct set of homologous alleles with very different activity often located close to boundaries of euchromatin and heterochromatin domains. We also identified a set of allele-biased topologically associating domains (TADs) that were small sized and had higher gene density. Thousands of genes show allelically biased expression (ABE) with false discovery rate < 0.05, and 98% of them have no allelic switching during G1, S, and G2/M phases. A subset of ABE genes are preferentially localized near TAD boundaries, enriched with chromatin organization transcription factor binding sites, and contained higher number of sequence variants in CCCTC-binding factor sites. Our results extend previous findings of sequence variation as a basis for unequal functional parental genomes. Investigation of haplotype-resolved form-function dynamics may further our understanding of phenotypic traits, genetic diseases, vulnerability to complex disorders, and the development of cancers.

genomics

Integrative pipeline for profiling DNA copy number and inferring tumor phylogeny

SummaryCopy number variation is an important and abundant source of variation in the human genome, which has been associated with a number of diseases, especially cancer. Massively parallel next-generation sequencing allows copy number profiling with fine resolution. Such efforts, however, have met with mixed successes, with setbacks arising partly from the lack of reliable analytical methods to meet the diverse and unique challenges arising from the myriad experimental designs and study goals in genetic studies. In cancer genomics, detection of somatic copy number changes and profiling of allele-specific copy number (ASCN) are complicated by experimental biases and artifacts as well as normal cell contamination and cancer subclone admixture. Furthermore, careful statistical modeling is warranted to reconstruct tumor phylogeny by both somatic ASCN changes and single nucleotide variants. Here we describe a flexible computational pipeline, MARATHON, which integrates multiple related statistical software for copy number profiling and downstream analyses in disease genetic studies.\n\nAvailability and implementationMARATHON is publicly available at https://github.com/yuchaojiang/MARATHON.\n\nContactyuchaoj@email.unc.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Spatiotemporal DNA Methylome Dynamics of the Developing Mammalian Fetus

Genetic studies have revealed an essential role for cytosine DNA methylation in mammalian development. However, its spatiotemporal distribution in the developing embryo remains obscure. Here, we profiled the methylome landscapes of 12 mouse tissues/organs at 8 developmental stages spanning from early embryogenesis to birth. Indepth analysis of these spatiotemporal epigenome maps systematically delineated ~2 million methylation variant regions and uncovered widespread methylation dynamics at nearly one-half million tissue-specific enhancers, whose human counterparts were enriched for variants involved in genetic diseases. Strikingly, these predicted regulatory elements predominantly lose CG methylation during fetal development, whereas the trend is reversed after birth. Accumulation of non-CG methylation within gene bodies of key developmental transcription factors coincided with their transcriptional repression during later stages of fetal development. These spatiotemporal epigenomic maps provide a valuable resource for studying gene regulation during mammalian tissue/organ progression and for pinpointing regulatory elements involved in human developmental diseases.

genomics

An Algorithm for Cellular Reprogramming

The day we understand the time evolution of subcellular elements at a level of detail comparable to physical systems governed by Newtons laws of motion seems far away. Even so, quantitative approaches to cellular dynamics add to our understanding of cell biology, providing data-guided frameworks that allow us to develop better predictions about, and methods for, control over specific biological processes and system-wide cell behavior. In this paper, we describe an approach to optimizing the use of transcription factors (TFs) in the context of cellular reprogramming. We construct an approximate model for the natural evolution of a cell cycle synchronized population of human fibroblasts, based on data obtained by sampling the expression of 22,083 genes at several time points along the cell cycle. In order to arrive at a model of moderate complexity, we cluster gene expression based on the division of the genome into topologically associating domains (TADs) and then model the dynamics of the TAD expression levels. Based on this dynamical model and known bioinformatics, such as transcription factor binding sites (TFBS) and functions, we develop a methodology for identifying the top transcription factor candidates for a specific cellular reprogramming task. The approach used is based on a device commonly used in optimal control. Our data-guided methodology identifies a number of transcription factors previously validated for reprogramming and/or natural differentiation. Our findings highlight the immense potential of dynamical models, mathematics, and data-guided methodologies for improving strategies for control over biological processes.\n\nSignificance StatementReprogramming the human genome toward any desirable state is within reach; application of select transcription factors drives cell types toward different lineages in many settings. We introduce the concept of data-guided control in building a universal algorithm for directly reprogramming any human cell type into any other type. Our algorithm is based on time series genome transcription and architecture data and known regulatory activities of transcription factors, with natural dimension reduction using genome architectural features. Our algorithm predicts known reprogramming factors, top candidates for new settings, and ideal timing for application of transcription factors. This framework can be used to develop strategies for tissue regeneration, cancer cell reprogramming, and control of dynamical systems beyond cell biology.

bioinformatics

A high-throughput assay to identify robust inhibitors of dynamin GTPase activity

Clathrin-mediated endocytosis is the major pathway by which cells internalize materials from the external environment. Dynamin, a large multidomain GTPase, is a key regulator of clathrin-mediated endocytosis. It assembles at the necks of invaginated clathrin-coated pits and, through GTP hydrolysis, catalyzes scission and release of clathrin-coated vesicles from the plasma membrane. Several small molecule inhibitors of dynamins GTPase activity, such as Dynasore and Dyngo-4a, are currently available, although their specificity has been brought into question. Previous screens for these inhibitors measured dynamins stimulated GTPase activity due to lack of sufficient sensitivity, hence the mechanisms by which they inhibit dynamin are uncertain. We report a highly sensitive fluorescence-based assay capable of detecting dynamins basal GTPase activity under conditions compatible with high throughput screening. Utilizing this optimized assay, we conducted a pilot screen of 8000 compounds and identified several \"hits\" that inhibit the basal GTPase activity of dynamin-1. Subsequent dose-response curves were used to validate the activity of these compounds. Interestingly, we found neither Dynasore nor Dyngo-4a inhibited dynamins basal GTPase activity, although both inhibit assembly-stimulated GTPase activity. This assay provides the basis for a more extensive search for robust dynamin inhibitors.

biochemistry

Genome Architecture Leads a Bifurcation in Cell Identity

Genome architecture is important in transcriptional regulation and study of its features is a critical part of fully understanding cell identity. Altering cell identity is possible through overexpression of transcription factors (TFs); for example, fibroblasts can be reprogrammed into muscle cells by introducing MYOD1. How TFs dynamically orchestrate genome architecture and transcription as a cell adopts a new identity during reprogramming is not well understood. Here we show that MYOD1-mediated reprogramming of human fibroblasts into the myogenic lineage undergoes a critical transition, which we refer to as a bifurcation point, where cell identity definitively changes. By integrating knowledge of genome-wide dynamical architecture and transcription, we found significant chromatin reorganization prior to transcriptional changes that marked activation of the myogenic program. We also found that the local architectural and transcriptional dynamics of endogenous MYOD1 and MYOG reflected the global genomic bifurcation event. These TFs additionally participate in entrainment of biological rhythms. Understanding the system-level genome dynamics underlying a cell fate decision is a step toward devising more sophisticated reprogramming strategies that could be used in cell therapies.

cell biology

Mapping Human Hematopoietic Hierarchy At Single Cell Resolution By Microwell-seq

The classical hematopoietic hierarchy, which is mainly built with fluorescence-activated cell sorting (FACS) technology, proves to be inaccurate in recent studies. Single cell RNA-seq (scRNA-seq) analysis provides a solution to overcome the limit of FACS-based cell type definition system for the dissection of complex cellular hierarchy. However, large-scale scRNA-seq is constrained by the throughput and cost of traditional methods. Here, we developed Microwell-seq, a high-throughput and low-cost scRNA-seq platform using extremely simple devices. Using Microwell-seq, we constructed a single-cell resolution transcriptome atlas of human hematopoietic differentiation hierarchy by profiling more than 50,000 single cells throughout adult human hematopoietic system. We found that adult human hematopoietic stem and progenitor cell (HSPC) compartment is dominated by progenitors primed with lineage specific regulators. Our analysis revealed differentiation pathways for each cell types, through which HSPCs directly progress to lineage biased progenitors before differentiation. We propose a revised adult human hematopoietic hierarchy independent of oligopotent progenitors. Our study also demonstrates the broad applicability of Microwell-seq technology.

cell biology

Fast functional annotation of metagenomic shotgun data by DNA alignment to a microbial gene catalog

BackgroundMetagenomic shotgun sequencing is becoming increasingly popular to study microbes associated with the human body and in environmental samples. A key goal of shotgun metagenomic sequencing is to identify gene functions and metabolic pathways that differ between samples or conditions. However, current methods to identify function in the large number of reads in a high-throughput sequence data file rely on the computationally intensive and low stringency approach of mapping each read to a generic database of proteins or reference microbial genomes.\n\nResultsWe have developed an alternative analysis approach for shotgun metagenomic sequence data utilizing Bowtie2 DNA-DNA alignment of the reads to a database of well annotated genes compiled from human microbiome data. This method is rapid, and provides high stringency matches (>90% DNA sequence identity) of shotgun metagenomics reads to genes with annotated functions. We demonstrate the use of this method with synthetic data, Human Microbiome Project shotgun metagenomic data sets, and data from a study of liver disease. Differentially abundant KEGG gene functions can be detected in these experiments.\n\nConclusionsFunctional annotation of metagenomic shotgun sequence reads can be accomplished by rapid DNA-DNA matching to a custom database of microbial sequences using the Bowtie2 sequence alignment tool. This method can be used for a variety of microbiome studies and allows functional analysis which is otherwise computationally demanding. This rapid annotation method is freely available as a Galaxy workflow within a Docker image.

bioinformatics

Applying Mondrian Cross-Conformal Prediction to Estimate Prediction Confidence on Large Imbalanced Bioactivity Datasets

Conformal prediction has been proposed as a more rigorous way to define prediction confidence compared to other application domain concepts that have earlier been used for QSAR modelling. One main advantage of such a method is that it provides a prediction region potentially with multiple predicted labels, which contrasts to the single valued (regression) or single label (classification) output predictions by standard QSAR modelling algorithms. Standard conformal prediction might not be suitable for imbalanced datasets. Therefore, Mondrian cross-conformal prediction (MCCP) which combines the Mondrian inductive conformal prediction with cross-fold calibration sets has been introduced. In this study, the MCCP method was applied to 18 publicly available datasets that have various imbalance levels varying from 1:10 to 1:1000 (ratio of active/inactive compounds). Our results show that MCCP in general performed well on cheminformatics datasets with various imbalance levels. More importantly, the method not only provides confidence of prediction and prediction regions compared to standard machine learning methods, but also produces valid predictions for the minority class. In addition, a compound similarity based nonconformity measure was investigated. Our results demonstrate that although it gives valid predictions, its efficiency is much worse than nonconformity measures obtained from supervised learning.

pharmacology and toxicology

Array-based sequencing of filaggrin gene for comprehensive detection of disease-associated variants

The filaggrin gene (FLG) is essential for skin differentiation and epidermal barrier formation with links to skin diseases, however it has a highly repetitive nucleotide sequence containing very limited stretches of unique nucleotides for precise mapping to reference genomes. Sequencing strategies using polymerase chain reaction (PCR) and conventional Sanger sequencing have been successful for complete FLG coding DNA sequence amplification to identify pathogenic mutations but this time-consuming, labour intensive method has restricted utility. Next-generation sequencing (NGS) offers obvious benefits to accelerate FLG analysis but standard re-sequencing techniques such as oligoprobe-based exome or customized targeted-capture can be expensive, especially for a single target gene of interest. We therefore designed a protocol to improve FLG sequencing throughput using a set of FLG-specific PCR primer assays compatible with microfluidic amplification, multiplexing and current NGS protocols. Using DNA reference samples with known FLG genotypes for benchmarking, this protocol is shown to be concordant for variant detection across different sequencing methodologies. We applied this methodology to analyze cohorts from ethnicities previously not studied for FLG variants and demonstrate usefulness for discovery projects. This comprehensive coverage sequencing protocol is labour-efficient and offers an affordable solution to scale up FLG sequencing for larger cohorts. Robust and rapid FLG sequencing can improve patient stratification for research projects and provide a framework for gene specific diagnosis in the future.

genetics

Direct visualization of transcriptional activation by physical enhancer-promoter proximity

A long-standing question in metazoan gene regulation is how remote enhancers communicate with their target promoters over long distances. Combining genome editing and quantitative live imaging we simultaneously visualize physical enhancer-promoter communication and transcription in Drosophila embryos. Enhancers regulating pair rule stripes of even-skipped expression activate transcription of a reporter gene over a distance of 150 kb. We show in individual cells that activation only occurs after the enhancer comes into close proximity with its regulatory target and that upon dissociation transcription ceases almost immediately. We further observe distinct topological conformations of the eve locus, depending on the spatial identity of the activating stripe enhancer. In addition, long-range activation results in transcriptional competition at the endogenous eve locus, causing corresponding developmental defects. Overall, we demonstrate that sustained physical proximity and enhancer-promoter engagement are required for enhancer action, and we provide a path to probe the implications of long-range regulation on cellular fates.

biophysics

Fast Assembling of Neuron Fragments in Serial 3D Sections

AbstractReconstructing neurons from 3D image-stacks of serial sections of thick brain tissue is very time-consuming and often becomes a bottleneck in high-throughput brain mapping projects. We developed NeuronStitcher, a software suite for stitching non-overlapping neuron fragments reconstructed in serial 3D image sections. With its efficient algorithm and user-friendly interface, NeuronStitcher has been used successfully to reconstruct very large and complex human and mouse neurons.

neuroscience

Efficient repositioning of approved drugs as anti-HIV agents using Anti-HIV-Predictor

Development of new, effective and affordable drugs against HIV is urgently needed. In this study, we developed a worlds first web server called Anti-HIV-Predictor (http://bsb.kiz.ac.cn:70/hivpre) for predicting anti-HIV activity of given compounds. This server is rapid and accurate (accuracy >93% and AUC > 0.958). We applied the server to screen 1835 approved drugs for anti-HIV therapy. Totally 67 drugs were predicted to have anti-HIV activity, 25 of which are anti-HIV drugs. Then we experimentally evaluated 35 predicted new anti-HIV compounds by assays of syncytia formation, p24 quantification, cytotoxicity. Finally, we repurposed 7 approved drugs (cetrorelix, dalbavancin, daunorubicin, doxorubicin, epirubicin, idarubicin and valrubicin) as new anti-HIV agents. The original indication of these drugs is involved in a variety of diseases such as female infertility and cancer. Anti-HIV-Predictor and the 7 repurposed anti-HIV agents provided here demonstrate the efficacy of this strategy for discovery of new anti-HIV agents.

bioinformatics