bioRxiv ScienceSearch

Biology subjects

Zhu, S.

Publications and source records attributed to Zhu, S..

16 recordsLinked to original sources

NetGO: Improving Large-scale Protein Function Prediction with Massive Network Information

Automated function prediction (AFP) of proteins is of great significance in biology. In essence, AFP is a large-scale multi-label classification over pairs of proteins and GO terms. Existing AFP approaches, however, have their limitations on both sides of proteins and GO terms. Using various sequence information and the robust learning to rank (LTR) framework, we have developed GOLabeler, a state-of-the-art approach of CAFA3, which overcomes the limitation of the GO term side, such as imbalanced GO terms. Unfortunately, for the protein side issue, available abundant protein information, except for sequences, have not been effectively used for large-scale AFP in CAFA. We propose NetGO that is able to improve large-scale AFP with massive network information. The novelties of NetGO have threefold in using network information: 1) the powerful LTR framework of NetGO efficiently and effectively integrates both sequence and network information, which can easily make large-scale AFP; 2) NetGO can use whole and massive network information of all species (>2000) in STRING (other than only high confidence links and/or some specific species); and 3) NetGO can still use network information to annotate a protein by homology transfer even if it is not covered in STRING. Under numerous experimental settings, we examined the performance of NetGO, such as general performance comparison, species-specific prediction, and prediction on difficult proteins, by using training and test data separated by time-delayed settings of CAFA. Experimental results have clearly demonstrated that NetGO outperforms GOLabeler, DeepGO, and other compared baseline methods significantly. In addition, several interesting findings from our experiments on NetGO would be useful for future AFP research.

bioinformatics

Epigenomic landscape of the human pathogen Clostridium difficile

Clostridioides difficile is a leading cause of health care-associated infections. Although significant progress has been made in the understanding of its genome, the epigenome of C. difficile and its functional impact has not been systematically explored. Here, we performed the first comprehensive DNA methylome analysis of C. difficile using 36 human isolates and observed great epigenomic diversity. We discovered an orphan DNA methyltransferase with a well-defined specificity whose corresponding gene is highly conserved across our dataset and in all ~300 global C. difficile genomes examined. Inactivation of the methyltransferase gene negatively impacted sporulation, a key step in C. difficile disease transmission, consistently supported by multi-omics data, genetic experiments, and a mouse colonization model. Further experimental and transcriptomic analysis also suggested that epigenetic regulation is associated with cell length, biofilm formation, and host colonization. These findings open up a new epigenetic dimension to characterize medically relevant biological processes in this critical pathogen. This work also provides a set of methods for comparative epigenomics and integrative analysis, which we expect to be broadly applicable to bacterial epigenomics studies.

genomics

Single-cell RNA-seq reveals dynamic transcriptome profiling in human early neural differentiation

BackgroundInvestigating cell fate decision and subpopulation specification in the context of the neural lineage is fundamental to understanding neurogenesis and neurodegenerative diseases. The differentiation process of neural-tube-like rosettes in vitro is representative of neural tube structures, which are composed of radially organized, columnar epithelial cells and give rise to functional neural cells. However, the underlying regulatory network of cell fate commitment during early neural differentiation remains elusive.\n\nResultsIn this study, we investigated the genome-wide transcriptome profile of single cells from six consecutive reprogramming and neural differentiation time points and identified cellular subpopulations present at each differentiation stage. Based on the inferred reconstructed trajectory and the characteristics of subpopulations contributing the most towards commitment to the central nervous system (CNS) lineage at each stage during differentiation, we identified putative novel transcription factors in regulating neural differentiation. In addition, we dissected the dynamics of chromatin accessibility at the neural differentiation stages and revealed active c/s-regulatory elements for transcription factors known to have a key role in neural differentiation as well as for those that we suggest are also involved. Further, communication network analysis demonstrated that cellular interactions most frequently occurred among embryoid body (EB) stage and each cell subpopulation possessed a distinctive spectrum of ligands and receptors associated with neural differentiation which could reflect the identity of each subpopulation.\n\nConclusionsOur study provides a comprehensive and integrative study of the transcriptomics and epigenetics of human early neural differentiation, which paves the way for a deeper understanding of the regulatory mechanisms driving the differentiation of the neural lineage.

developmental biology

Demographic inference using particle filters for continuous Markov jump processes

Demographic events shape a populations genetic diversity, a process described by the coalescent-with-recombination (CwR) model that relates demography and genetics by an unobserved sequence of genealogies. The space of genealogies over genomes is large and complex, making inference under this model challenging. We approximate the CwR with a continuous-time and -space Markov jump process. We develop a particle filter for such processes, using way-points to reduce the problem to the discrete-time case, and generalising the Auxiliary Particle Filter for discrete-time models. We use Variational Bayes for parameter inference to model the uncertainty in parameter estimates for rare events, avoiding biases seen with Expectation Maximization. Using real and simulated genomes, we show that past population sizes can be accurately inferred over a larger range of epochs than was previously possible, opening the possibility of jointly analyzing multiple genomes under complex demographic models. Code is available at https://github.com/luntergroup/smcsmc MSC 2010 subject classificationsPrimary 60G55, 62M05, 62M20, 62F15; secondary 92D25.

evolutionary biology

Comprehensive analysis of immune evasion in breast cancer by single-cell RNA-seq

The tumor microenvironment is composed of numerous cell types, including tumor, immune and stromal cells. Cancer cells interact with the tumor microenvironment to suppress anticancer immunity. In this study, we molecularly dissected the tumor microenvironment of breast cancer by single-cell RNA-seq. We profiled the breast cancer tumor microenvironment by analyzing the single-cell transcriptomes of 52,163 cells from the tumor tissues of 15 breast cancer patients. The tumor cells and immune cells from individual patients were analyzed simultaneously at the single-cell level. This study explores the diversity of the cell types in the tumor microenvironment and provides information on the mechanisms of escape from clearance by immune cells in breast cancer.\n\nOne Sentence SummaryLandscape of tumor cells and immune cells in breast cancer by single cell RNA-seq

cancer biology

The Vibrio H-ring facilitates the outer membrane penetration of polar-sheathed flagellum

The bacterial flagellum has evolved as one of the most remarkable nanomachines in nature. It provides swimming and swarming motilities that are often essential for the bacterial life cycle and for pathogenesis. Many bacteria such as Salmonella and Vibrio species use flagella as an external propeller to move to favorable environments, while spirochetes utilize internal periplasmic flagella to drive a serpentine movement of the cell bodies through tissues. Here we use cryo-electron tomography to visualize the polar-sheathed flagellum of Vibrio alginolyticus with particular focus on a Vibrio specific feature, the H-ring. We characterized the H-ring by identifying its two components FlgT and FlgO. Surprisingly, we discovered that the majority of flagella are located within the periplasmic space in the absence of the H-ring, which are dramatically different from external flagella in wild-type cells. Our results indicate the H-ring has a novel function in facilitating the penetration of the outer membrane and the assembly of the external sheathed flagella. This unexpected finding is however consistent with the notion that the flagella have evolved to adapt highly diverse needs by receiving or removing accessary genes.\n\nSignificance StatementFlagellum is the major organelle for motility in many bacterial species. While most bacteria possess external flagella such as the multiple peritrichous flagella found in Escherichia coli and Salmonella enterica or the single polar-sheathed flagellum in Vibrio spp., spirochetes uniquely assemble periplasmic flagella, which are embedded between their inner and outer membranes. Here, we show for the first time that the external flagella in Vibrio alginolyticus can be changed as periplasmic flagella by deleting two flagellar genes. The discovery here may provide a new paradigm to understand the molecular basis underlying flagella assembly, diversity, and evolution.

microbiology

Recovered and dead outcome patients caused by influenza A (H7N9) virus infection show different pro-inflammatory cytokine dynamics during disease progress and its application in real-time prognosis

The persistent circulation of influenza A(H7N9) virus within poultry markets and human society leads to sporadic epidemics of influenza infections. Severe pneumonia and acute respiratory distress syndrome (ARDS) caused by the virus lead to high morbidity and mortality rates in patients. Hyper induction of pro-inflammatory cytokines, which is known as \"cytokine storm\", is closely related to the process of viral infection. However, systemic analyses of H7N9 induced cytokine storm and its relationship with disease progress need further illuminated. In our study we collected 75 samples from 24 clinically confirmed H7N9-infected patients at different time points after hospitalization. Those samples were divided into three groups, which were mild, severe and fatal groups, according to disease severity and final outcome. Human cytokine antibody array was performed to demonstrate the dynamic profile of 80 cytokines and chemokines. By comparison among different prognosis groups and time series, we provide a more comprehensive insight into the hypercytokinemia caused by H7N9 influenza virus infection. Different dynamic changes of cytokines/chemokines were observed in H7N9 infected patients with different severity. Further, 33 cytokines or chemokines were found to be correlated with disease development and 11 of them were identified as potential therapeutic targets. Immuno-modulate the cytokine levels of IL-8, IL-10, BLC, MIP-3a, MCP-1, HGF, OPG, OPN, ENA-78, MDC and TGF-{beta} 3 are supposed to be beneficial in curing H7N9 infected patients. Apart from the identification of 35 independent predictors for H7N9 prognosis, we further established a real-time prediction model with multi-cytokine factors for the first time based on maximal relevance minimal redundancy method, and this model was proved to be powerful in predicting whether the H7N9 infection was severe or fatal. It exhibited promising application in prognosing the outcome of a H7N9 infected patients and thus help doctors take effective treatment strategies accordingly.

immunology

Characteristics and origins of non-functional Pm21 alleles in Dasypyrum villosum and wheat genetic stocks

Most Dasypyrum villosum resources are highly resistant to wheat powdery mildew that carries Pm21 alleles. However, in the previous studies, four D. villosum lines (DvSus-1 [~] DvSus-4) and two wheat-D. villosum addition lines (DA6V#1 and DA6V#3) were reported to be susceptible to powdery mildew. In the present study, the characteristics of non-functional Pm21 alleles in the above resources were analyzed after Sanger sequencing. The results showed that loss-of-functions of Pm21 alleles Pm21-NF1 [~] Pm21-NF3 isolated from DvSus-1, DvSus-2/DvSus-3 and DvSus-4 were caused by two potential point mutations, a 1-bp deletion and a 1281-bp insertion, respectively. The non-functional Pm21 alleles in DA6V#1 and DA6V#3 were same to that in DvSus-4 and DvSus-2/DvSus-3, respectively, indicating that the susceptibilities of the two wheat genetic stocks came from their D. villosum donors. The origins of non-functional Pm21 alleles were also investigated in this study. Except the target variants involved, the sequences of Pm21-NF2 and Pm21-NF3 were identical to that of Pm21-F2 and Pm21-F3 in the resistant D. villosum lines DvRes-2 and DvRes-3, derived from the accessions GRA961 and GRA1114, respectively. It was suggested that the non-functional alleles Pm21-NF2 and Pm21-NF3 originated from the wild-type alleles Pm21-F2 and Pm21-F3. In summary, this study gives an insight into the sequence characteristics of non-functional Pm21 alleles and their origins in natural population of D. villosum.

plant biology

Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity

Integrative analysis of multi-omics layers at single cell level is critical for accurate dissection of cell-to-cell variation within certain cell populations. Here we report scCAT-seq, a technique for simultaneously assaying chromatin accessibility and the transcriptome within the same single cell. We show that the combined single cell signatures enable accurate construction of regulatory relationships between cis-regulatory elements and the target genes at single-cell resolution, providing a new dimension of features that helps direct discovery of regulatory patterns specific to distinct cell identities. Moreover, we generated the first single cell integrated maps of chromatin accessibility and transcriptome in human pre-implantation embryos and demonstrated the robustness of scCAT-seq in the precise dissection of master transcription factors in cells of distinct states during embryo development. The ability to obtain these two layers of omics data will help provide more accurate definitions of \"single cell state\" and enable the deconvolution of regulatory heterogeneity from complex cell populations.

genomics

Rotational 3D mechanogenomic Turing patterns of human colon Caco-2 cells during differentiation

Recent reports suggest that actomyosin meshwork act in a mechanobiological manner alter cell/nucleus/tissue morphology, including human colon epithelial Caco-2 cancer cells that form polarized 2D epithelium or 3D sphere/tube when placed in different culture conditions. We observed the rotational motion of the nucleus in Caco-2 cells in vitro that appears to be driven by actomyosin network prior to the formation of a differentiated confluent epithelium. Caco-2 cell monolayer preparations demonstrated 2D patterns consistent with Allan Turings \"gene morphogen\" hypothesis based on live cell imaging analysis of apical tight junctions indicating the actomyosin meshwork. Caco-2 cells in 3D culture are frequently used as a model to study 3D epithelial morphogenesis involving symmetric and asymmetric cell divisions. Differentiation of Caco-2 cells in vitro demonstrated similarity to intestinal enterocyte differentiation along the human colon crypt axis. We observed rotational 3D patterns consistent with gene morphogens during Caco-2 cell differentiation. Single- to multi-cell ring/torus-shaped genomes were observed that were similar to complex fractal Turing patterns extending from a rotating torus centre in a spiral pattern consistent with gene morphogen motif. Rotational features of the epithelial cells may contribute to well-described differentiation from stem cells to the luminal colon epithelium along the crypt axis. This dataset may be useful to study the role of mechanobiological processes and the underlying molecular mechanisms as determinants of cellular and tissue architecture in space and time, which is the focal point of the 4D nucleome initiative.

bioinformatics

Early mannitol-triggered changes in the Arabidopsis leaf (phospho)proteome

Drought is one of the most detrimental environmental stresses to which plants are exposed. Especially mild drought is relevant to agriculture and significantly affects plant growth and development. In plant research, mannitol is often used to mimic drought stress and study the underlying responses. In growing leaf tissue of plants exposed to mannitol-induced stress, a highly-interconnected gene regulatory network is induced. However, early signaling and associated protein phosphorylation events that likely precede part of these transcriptional changes are largely unknown. Here, we performed a full proteome and phosphoproteome analysis on growing leaf tissue of Arabidopsis plants exposed to mild mannitol-induced stress and captured the fast (within the first half hour) events associated with this stress. Based on this in-depth data analysis, 167 and 172 differentially regulated proteins and phosphorylated sites were found back, respectively. Additionally, we identified H(+)-ATPASE 2 (AHA2) and CYSTEINE-RICH REPEAT SECRETORY PROTEIN 38 (CRRSP38) as novel regulators of shoot growth under osmotic stress.\n\nHighlightWe captured early changes in the Arabidopsis leaf proteome and phosphoproteome upon mild mannitol stress and identified AHA2 and CRRSP38 as novel regulators of shoot growth under osmotic stress

plant biology

The small molecule KHS101 induces bioenergetic dysfunction in glioblastoma cells through inhibition of mitochondrial HSPD1

Pharmacological inhibition of uncontrolled cell growth with small molecule inhibitors is a potential strategy against glioblastoma multiforme (GBM), the most malignant primary brain cancer. Phenotypic profiling of the neurogenic small molecule KHS101 revealed the chemical induction of lethal cellular degradation in molecularly-diverse GBM cells, independent of their tumor subtype, whereas non-cancerous brain cells remained viable. Mechanism-of-action (MOA) studies showed that KHS101 specifically bound and inhibited the mitochondrial chaperone HSPD1. In GBM but not non-cancerous brain cells, KHS101 elicited the aggregation of an enzymatic network that regulates energy metabolism. Compromised glycolysis and oxidative phosphorylation (OXPHOS) resulted in the metabolic energy depletion in KHS101-treated GBM cells. Consistently, KHS101 induced key mitochondrial unfolded protein response factor DDIT3 in vitro and in vivo, and significantly reduced intracranial GBM xenograft tumor growth upon systemic administration, without discernible side effects. These findings suggest targeting of HSPD1-dependent oncometabolic pathways as an anti-GBM therapy.

cancer biology

Map-based cloning of the gene Pm21 that confers broad spectrum resistance to wheat powdery mildew

Common wheat (Triticum aestivum L.) is one of the most important cereal crops. Wheat powdery mildew caused by Blumeria graminis f. sp. tritici (Bgt) is a continuing threat to wheat production. The Pm21 gene, originating from Dasypyrum villosum, confers high resistance to all known Bgt races and has been widely applied in wheat breeding in China. In this research, we identify Pm21 as a typical coiled-coil, nucleotide-binding site, leucine-rich repeat gene by an integrated strategy of resistance gene analog (RGA)-based cloning via comparative genomics, physical and genetic mapping, BSMV-induced gene silencing (BSMV-VIGS), large-scale mutagenesis and genetic transformation.

genetics

N6-Methyladenine DNA Modification in Human Genome

DNA N6-methyladenine (6mA) modification is the most prevalent DNA modification in prokaryotes, but whether it exists in human cells and whether it plays a role in human diseases remain enigmatic. Here, we showed that 6mA is extensively present in human genome, and we cataloged 881,240 6mA sites accounting for [~]0.051% of the total adenines. [G/C]AGG[C/T] was the most significantly associated motif with 6mA modification. 6mA sites were enriched in the coding regions and mark actively transcribed genes in human cells. We further found that DNA N6-methyladenine and N6-demethyladenine modification in human genome were mediated by methyltransferase N6AMT1 and demethylase ALKBH1, respectively. The abundance of 6mA was significantly lower in cancers, accompaning with decreased N6AMT1 and increased ALKBH1 levels, and down-regulation of 6mA modification levels promoted tumorigenesis. Collectively, our results demonstrate that DNA 6mA modification is extensively present in human cells and the decrease of genomic DNA 6mA promotes human tumorigenesis.

genomics

GOLabeler: Improving Sequence-based Large-scale Protein Function Prediction by Learning to Rank

Motivation: Gene Ontology (GO) has been widely used to annotate functions of proteins and understand their biological roles. Currently only {inverted exclamation}1% of more than 70 million proteins in UniProtKB have experimental GO annotations, implying the strong necessity of automated function prediction (AFP) of proteins, where AFP is a hard multi-label classification problem due to one protein with a diverse number of GO terms. Most of these proteins have only sequences as input information, indicating the importance of sequence-based AFP (SAFP: sequences are the only input). Furthermore, homology-based SAFP tools are competitive in AFP competitions, while they do not necessarily work well for so-called difficult proteins, which have {inverted exclamation}60% sequence identity to proteins with annotations already. Thus, the vital and challenging problem now is to develop a method for SAFP, particularly for difficult proteins.\n\nMethods: The key of this method is to extract not only homology information but also diverse, deep-rooted information/evidence from sequence inputs and integrate them into a predictor in an efficient and also effective manner. We propose GOLabeler, which integrates five component classifiers, trained from different features, including GO term frequency, sequence alignment, amino acid trigram, domains and motifs, and biophysical properties, etc., in the framework of learning to rank (LTR), a new paradigm of machine learning, especially powerful for multi-label classification.\n\nResults: The empirical results obtained by examining GOLabeler extensively and thoroughly by using large-scale datasets revealed numerous favorable aspects of GOLabeler, including significant performance advantage over state-of-the-art AFP methods.\n\nContact: zhusf@fudan.edu.cn

bioinformatics

Characterization Of Imprinted Genes In Rice Reveals Post-Fertilization Regulation And Conservation At Some Loci Of Imprinting In Plant Species

Genomic imprinting is an epigenetic phenomenon by which certain genes display monoallelic expression in a parent-of-origin-dependent manner. Hundreds of imprinted genes have been identified from several plant species. Here we identified, with a high level of confidence, 208 imprinted candidates from rice. Imprinted genes of rice showed limited association to the transposable elements, which is contrast to the findings in Arabidopsis. Generally, imprinting of rice is conserved within species, but intraspecific variations were confirmed here. Imprinting between cultivated rice and wild rice are likely similar. The imprinted genes of rice do not show significant selective signatures overall, which suggests that domestication imposes limited evolutionary effects on genomic imprinting of rice. Though the conservation of imprinting in plants is limited, here we prove that some loci tend to be imprinted in different species. In addition, our results suggest that differential epigenetic regulation between parental alleles can be established either prior to or post-fertilization. The imprinted 24-nt small RNAs, but not the 21-nt ones, likely involve the regulation of imprinting in an opposite parental-allele targeting manner. Together, our findings suggest that regulation of imprinting can be very diverse, and genomic imprinting as well as imprinted genes have essential evolutionary and biological significance.

plant biology