bioRxiv ScienceSearch

Biology subjects

Zhang, B.

Publications and source records attributed to Zhang, B..

At least 37 records · Page 2Linked to original sources

Epitope-based vaccine design yields fusion peptide-directed antibodies that neutralize diverse strains of HIV-1

A central goal of HIV-1-vaccine research is the elicitation of antibodies capable of neutralizing diverse primary isolates of HIV-1. Here we show that focusing the immune response to exposed N-terminal residues of the fusion peptide, a critical component of the viral entry machinery and the epitope of antibodies elicited by HIV-1 infection, through immunization with fusion peptide-coupled carriers and prefusion-stabilized envelope trimers, induces cross-clade neutralizing responses. In mice, these immunogens elicited monoclonal antibodies capable of neutralizing up to 31% of a cross-clade panel of 208 HIV-1 strains. Crystal and cryo-electron microscopy structures of these antibodies revealed fusion peptide-conformational diversity as a molecular explanation for the cross-clade neutralization. Immunization of guinea pigs and rhesus macaques induced similarly broad fusion peptide-directed neutralizing responses suggesting translatability. The N terminus of the HIV-1-fusion peptide is thus a promising target of vaccine efforts aimed at eliciting broadly neutralizing antibodies.

immunology

Widespread and polymorphous noncoding amino acid residues in human sperm proteome

Proteins are usually deciphered by translation of the coding genome; however, their amino acid residues are seldom determined directly across the proteome. Herein, we describe a systematic workflow for identifying all possible protein residues that differ from the coding genome, termed noncoded amino acids (ncAAs). By measuring the mass differences between the coding amino acids and the actual protein residues in human spermatozoa, over a million nonzero delta masses were detected, fallen into 424 high-quality Gaussian clusters and 571 high-confidence ncAAs spanning 29,053 protein sites. Most ncAAs are novel with unresolved side-chains and discriminative between healthy individuals and patients with oligoasthenospermia. For validation, 40 out of 98 ncAAs that matched with amino acid substitutions were confirmed by exon sequencing. This workflow revealed the widespread existence of previously unreported ncAAs in the sperm proteome, which represents a new dimension on the understanding of amino acid polymorphisms at the proteomic level.\n\nHighlightsO_LI571 ncAAs spanning 108,000 protein sites were identified in human sperm proteome.\nC_LIO_LIMost ncAAs are novel with unresolved sidechains and found at unreported protein sites.\nC_LIO_LIExon sequencing confirmed 40 of 98 ncAAs that matched with amino acid substitutions.\nC_LIO_LIMany ncAAs are linked with disease and have potential for diagnosis and targeting.\nC_LI\n\neTOC BlurbWe describe a systematic identification of all possible protein residues that were not encoded by their genomic sequences. A total of 571 high-confidence most novel noncoded amino acids were identified in human sperm proteome, corresponding to over 108,000 ncAA-containing protein sites. For validation, 40 out of 98 ncAAs that matched to amino acid substitutions were confirmed by exon sequencing. These ncAAs are discriminative between individuals and expand our understanding of amino acid polymorphisms in human proteomes and diseases.

molecular biology

Predicting three-dimensional genome organization with chromatin states

We introduce a computational model to simulate chromatin structure and dynamics. Starting from one-dimensional genomics and epigenomics data that are available for hundreds of cell types, this model enables de novo prediction of chromatin structures at five-kilo-base resolution. Simulated chromatin structures recapitulate known features of genome organization, including the formation of chromatin loops, topologically associating domains (TADs) and compartments, and are in quantitative agreement with chromosome conformation capture experiments and super-resolution microscopy measurements. Detailed characterization of the predicted structural ensemble reveals the dynamical flexibility of chromatin loops and the presence of cross-talk among neighboring TADs. Analysis of the models energy function uncovers distinct mechanisms for chromatin folding at various length scales.

biophysics

Gut microbiota density influences host physiology and is shaped by host and microbial factors

To identify factors that regulate gut microbiota density and the impact of varied microbiota density on health, we assayed this fundamental ecosystem property in fecal samples across mammals, human disease, and therapeutic interventions. Physiologic features of the host (carrying capacity) and the fitness of the gut microbiota shape microbiota density. Therapeutic manipulation of microbiota density in mice altered host metabolic and immune homeostasis. In humans, gut microbiota density was reduced in Crohns disease, ulcerative colitis, and ileal pouch-anal anastomosis. The gut microbiota in recurrent Clostridium difficile infection had lower density and reduced fitness that were restored by fecal microbiota transplantation. Understanding the interplay between microbiota and disease in terms of microbiota density, host carrying capacity, and microbiota fitness provide new insights into microbiome structure and microbiome targeted therapeutics.

microbiology

Dynamic virulence-related regions of the fungal plant pathogen Verticillium dahliae display remarkably enhanced sequence conservation

Selection pressure impacts genomes unevenly, as different genes adapt with differential speed to establish an organisms optimal fitness. Plant pathogens co-evolve with their hosts, which implies continuously adaption to evade host immunity. Effectors are secreted proteins that mediate immunity evasion, but may also typically become recognized by host immune receptors. To facilitate effector repertoire alterations, in many pathogens, effector genes reside in dynamic genomic regions that are thought to display accelerated evolution, a phenomenon that is captured by the two-speed genome hypothesis. The genome of the vascular wilt pathogen Verticillium dahliae has been proposed to obey to a similar two-speed regime with dynamic, lineage-specific regions that are characterized by genomic rearrangements, increased transposable element activity and enrichment in in planta-induced effector genes. However, little is known of the origin of, and sequence diversification within, these lineage-specific regions. Based on comparative genomics among Verticillium spp. we now show differential sequence divergence between core and lineage-specific genomic regions of V. dahliae. Surprisingly, we observed that lineage-specific regions display markedly increased sequence conservation. Since single nucleotide diversity is reduced in these regions, host adaptation seems to be merely achieved through presence/absence polymorphisms. Increased sequence conservation of genomic regions important for pathogenicity is an unprecedented finding for filamentous plant pathogens and signifies the diversity of genomic dynamics in host-pathogen co-evolution.

microbiology

Genome-wide identification of CDC34 that stabilizes EGFR and promotes lung carcinogenesis

To systematically identify ubiquitin pathway genes that are critical to lung carcinogenesis, we used a genome-wide silencing method in this study to knockdown 696 genes in non-small cell lung cancer (NSCLC) cells. We identified 31 candidates that were required for cell proliferation in two NSCLC lines, among which the E2 ubiquitin conjugase CDC34 represented the most significant one. CDC34 was elevated in tumor tissues in 67 of 102 (65.7%) NSCLCs, and smokers had higher CDC34 than nonsmokers. The expression of CDC34 was inversely associated with overall survival of the patients. Forced expression of CDC34 promoted, whereas knockdown of CDC34 inhibited lung cancer in vitro and in vivo. CDC34 bound EGFR and competed with E3 ligase c-Cbl to inhibit the polyubiquitination and subsequent degradation of EGFR. In EGFR-L858R and EGFR-T790M/Del(exon 19)-driven lung cancer in mice, knockdown of CDC34 by lentivirus mediated transfection of short hairpin RNA significantly inhibited tumor formation. These results demonstrate that an E2 enzyme is capable of competing with E3 ligase to inhibit ubiquitination and subsequent degradation of oncoprotein substrate, and CDC34 represents an attractive therapeutic target for NSCLCs with or without drug-resistant EGFR mutations.

cancer biology

CRISPR-typing PCR (ctPCR), a new Cas9-based DNA detection method

This study develops a new method for detecting and typing target DNA based on Cas9 nuclease, which was named as ctPCR, representing Cas9/sgRNA- or CRISPR-typing PCR. The technique can detect and discriminate target DNA easily, rapidly, specifically, and sensitively. This technique detects target DNA in three steps: (1) amplifying target DNA with PCR by using a pair of universal primers (PCR1); (2) treating PCR1 products with a process referred to as CAT, representing Cas9 cutting, A tailing and T adaptor ligation; (3) amplifying the CAT-treated DNA with PCR by using a pair of general-specific primers (gs-primers) (PCR2). The technique was verified by detecting HPV16 and HPV18 L1 gene in 13 different high-risk human papillomavirus (HPV) subtypes. The technique was also detected two high-risk HPVs (HPV16 and HPV18) in cervical carcinoma cells (HeLa and SiHa) by detecting the L1 and E6/E7 genes, respectively. In this method, PCR1 was performed to determine if the detected DNA sample contained the target DNA (such as virus infection), while PCR2 was performed to discriminate which genotypic target DNA was present in the detected DNA sample (such as virus subtypes). With these proof-of-concept experiments, this study provides a new CRISPR-based DNA detection and typing method.

biochemistry

The Function of the COPII Gene Paralogs SEC23A and SEC23B Are Interchangeable In Vivo

SEC23 is a core component of the coat protein-complex II (COPII)-coated vesicle, which mediates transport of secretory proteins from the endoplasmic reticulum (ER) to the Golgi1-3. Mammals express 2 paralogs for SEC23 (SEC23A and SEC23B). Though the SEC23 gene duplication dates back >500 million years, both SEC23s are ~85% identical at the amino acid sequence level. In humans, deficiency for SEC23A or SEC23B results in cranio-lenticulo-sutural dysplasia4 or congenital dyserythropoietic anemia type II (CDAII), respectively5. The disparate human syndromes and reports of secretory cargos with apparent paralog-specific dependence6,7, suggest unique functions for the two SEC23 paralogs. Here we show indistinguishable intracellular interactomes for human SEC23A and SEC23B, complementation of yeast SEC23 by both human and murine SEC23A/B paralogs, and the rescue of lethality resulting from Sec23b disruption in zebrafish by a Sec23a-expressing transgene. Finally, we demonstrate that the Sec23a coding sequence inserted into the endogenous murine Sec23b locus fully rescues the mortality and severe pancreatic phenotype previously reported with SEC23B-deficiency in the mouse8-10. Taken together, these data indicate that the disparate phenotypes of SEC23A and SEC23B deficiency likely result from evolutionary shifts in gene expression program rather than differences in protein function, a paradigm likely applicable to other sets of paralogous genes. These findings also suggest the potential for increased expression of SEC23A as a novel therapeutic approach to the treatment of CDAII, with potential relevance to other disorders due to mutations in paralogous genes.

evolutionary biology

Novel Compensatory Mechanisms Enable the Mutant KCNT1 Channels to Induce Seizures

Mutations in the sodium-activated potassium channel (KCNT1) gene are linked to epilepsy. Surprisingly, all KCNT1 mutations examined to date increase K+ current amplitude. These findings present a major neurophysiological paradox: how do gain-of-function KCNT1 mutations expected to silence neurons cause epilepsy? Here, we use Drosophila to show that expressing mutant KCNT1 in GABAergic neurons leads to seizures, consistent with the notion that silencing inhibitory neurons tips the balance towards hyperexcitation. Unexpectedly, mutant KCNT1 expressed in motoneurons also causes seizures. One striking observation is that mutant KCNT1 causes abnormally large and spontaneous EJPs (sEJPs). Our data suggest that these sEJPs result from local depolarization of synaptic terminals due to a reduction in Shaker channel levels and more active Na+ channels. Hence, we provide the first in vivo evidence that both disinhibition of inhibitory neurons and compensatory plasticity in motoneurons can account for the paradoxical effects of gain-of-function mutant KCNT1 in epilepsy.

neuroscience

Effects of regional differences on the urinary proteomes of healthy Chinese individuals

Urine is a promising biomarker source for clinical proteomics studies. Although regional physiological differences are common in multi-center clinical studies, the presence of significant differences in the urinary proteomes of individuals from different regions remains unknown. In this study, morning urine samples were collected from healthy urban residents in three regions of China and urinary proteins were preserved using a membrane-based method (Urimem). The urine proteomes of 27 normal samples were analyzed using LC-MS/MS and compared among the three regions. We identified 1,898 proteins from Urimem samples using label-free proteome quantification, of which 62 urine proteins were differentially expressed among the three regions. Hierarchical clustering analysis showed that inter-regional differences caused less significant changes in the urine proteome than inter-sex differences. Of the 62 differentially expressed proteins, 10 have been reported to be disease biomarkers in previous clinical studies. Urimem facilitates urinary protein storage for large-scale urine sample collection, and thus accelerates biobank development and urine biomarker studies employing proteomics approaches. Regional differences are a confounding factor influencing the urine proteome and should be considered in future multi-center biomarker studies.

physiology

canvasXpress: A versatile interactive high-resolution scientific multi-panel visualization toolkit

To the Editor: CanvasXpress (https://canvasxpress.org) was developed as the core visualization component for bioinformatics and systems biology analysis at Bristol-Myers Squibb and further enhanced by scientists around the world and served as a key visualization engine for many popular bioinformatics tools1,2,3,4,5,6. It offers a rich set of interactive plots to display scientific and genomics data, such as oncoprint of cancer mutations, heatmap, 3D scatter, violin, radar, and profile plots (Figure 1, canvasXpress plots arranged by canvasDesigner https://baohongz.github.io/canvasDesigner). Recently, the reproducibility and usability of the package in real world bioinformatics and clinical use cases have been improved significantly witnessed by continuous add-on features and wide adoption of the toolkit in the scientific communities. Furthermore, It is the first noteworthy package harmonizing real time interactive exploring and analyzing of big data, full-fledged customizing of look-n-feel, and producing multi-panel publication-ready figures in PDF format simultaneously.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=84 SRC=\"FIGDIR/small/186213_fig1.gif\" ALT=\"Figure 1\">\nView larger version (36K):\norg.highwire.dtl.DTLVardef@1d8392borg.highwire.dtl.DTLVardef@916a49org.highwire.dtl.DTLVardef@d90b05org.highwire.dtl.DTLVardef@162a8fc_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Versatile plots generated by canvasXpress. A) Oncoprint of cancer mutations; B) Gene expression heatmap with comprehensive annotations; C) 3D scatter plot; D) Violin plot; E) Radar plot with annotations; F) Profile plot of gene expression.\n\nC_FIG

bioinformatics

The effect of sonic hedgehog on motor neuron positioning in the spinal cord during chicken embryo development

Sonic hedgehog (Shh) is a vertebrate homologue of the secreted Drosophila protein hedgehog, and is expressed by the notochord and the floor plate in the developing spinal cord. Shh provides signals relevant for positional information, cell proliferation, and possibly cell survival depending on the time and location of the expression. Although the role of Shh in providing positional information in the neural tube has been experimentally proven, the exact underlying mechanism still remains unclear. In this study, we report that overexpression of Shh affects motor neuron positioning in the spinal cord during chicken embryo development by inducing abnormalities in the structure of the motor column and motor neuron integration. In addition, Shh overexpression inhibits the expression of dorsal transcription factors and commissural axon projections. Our results indicate that correct location of Shh expression is the key to the formation of the motor column. In conclusion, the overexpression of Shh in the spinal cord not only affects the positioning of motor neurons, but also induces abnormalities in the structure of the motor column.

developmental biology

Integrative analyses of splicing in the aging brain: role in susceptibility to Alzheimer’s Disease

We use deep sequencing to identify sources of variation in mRNA splicing in the dorsolateral prefrontal cortex (DLFPC) of 450 subjects from two prospective cohort studies of aging. Hundreds of aberrant pre-mRNA splicing events are reproducibly associated with Alzheimers Disease (AD). We also generate a catalog of splicing quantitative trait loci (sQTL) effects in the human cortex: splicing of 3,198 genes is influenced by genetic variation. sQTLs are enriched among those variants influencing DNA methylation and histone acetylation. In assessing known AD loci, we report that altered splicing is the mechanism for the effects of the PICALM, CLU, and PTK2B susceptibility alleles. Further, we leverage our sQTL catalog to identify genes whose aberrant splicing is associated with AD and mediated by genetics. This transcriptome-wide association study identified 21 genes with significant associations, many of which are found in AD GWAS loci, but 8 are in novel AD loci, including FUS, which is a known amyotrophic lateral sclerosis (ALS) gene. This highlights an intriguing shared genetic architecture that is further elaborated by the convergence of old and new AD genes in autophagy-lysosomal-related pathways already implicated in AD and other neurodegenerative diseases. Overall, this study of the aging brains transcriptome provides evidence that dysregulation of mRNA splicing is a feature of AD and is, in some genetically-driven cases, causal.

genomics

The Correlation between MicroRNA-199a and White Adipose Tissue in C57/BL6J Mice with High-Fat Diet

Understanding is emerging about microRNAs as mediators in the regulation of white adipose tissue (WAT) and obesity. The expression level of miR-199a in mice was investigated to test our hypothesis: miR-199a might be related to fat accumulation and try to find its target mRNA, which perhaps could propose strategies with a therapeutic potential affecting the fat storage. C57/BL6J mice were randomly assigned to either a control group or an obesity model group (n=10 in both groups). Control mice were fed a normal diet (fat: 10 kcal %) ad libitum for 12 weeks, and model mice were fed a high-fat diet (fat: 30 kcal %) ad libitum for 12 weeks to induce obesity. At the end of the experiment, body fat mass and the free fatty acids (FFAs) and triglycerides (TGs) in WAT were tested. Fat cell size was measured by hematoxylin-eosin (H&E) staining method. The fat mass of the model group was higher than that of the control group (P<0.05). In addition, the concentrations of the FFAs and TGs were higher (P<0.05) and the adipocyte count was lower (P<0.05) in the model group. We tested the expression levels of miR-199a and key adipogenic transcription factors, including peroxisome proliferator activated receptor gamma2 (PPAR{gamma}2), CCAAT/enhancer binding proteins alpha (C/EBP), adipocyte fatty acid-binding protein (aP2), and sterol regulatory element binding protein-1c (SREBP-1c). Up-regulated expression of miR-199a was observed in model group. Increased levels of miR-199a was accompanied by high expression levels of SREBP-1c. We found that the 3-UTR of SREBP-1c mRNA has a predicted binding site for miR-199a. Based on the current discoveries, we suggest that miR-199a may exert its action by binding to its target mRNA and cooperate with SREBP-1c to induce obesity. Therefore, if the predicted binding site is confirmed by further research, miR-199a may have therapeutic potential for obesity.\n\nAbbreviationsWAT, white adipose tissue; PPAR{gamma}2, peroxisome proliferator, activated receptor {gamma}2; C/EBP CCAAT/enhancer binding proteins ; aP2, adipocyte fatty acid-binding protein; SREBP-1c, sterol regulatory element binding protein-1c; HFD, high-fat diet.

molecular biology

Systematic mapping of chromatin state landscapes during mouse development

Embryogenesis requires epigenetic information that allows each cell to respond appropriately to developmental cues. Histone modifications are core components of a cells epigenome, giving rise to chromatin states that modulate genome function. Here, we systematically profile histone modifications in a diverse panel of mouse tissues at 8 developmental stages from 10.5 days post conception until birth, performing a total of 1,128 ChIP-seq assays across 72 distinct tissue-stages. We combine these histone modification profiles into a unified set of chromatin state annotations, and track their activity across developmental time and space. Through integrative analysis we identify dynamic enhancers, reveal key transcriptional regulators, and characterize the role of chromatin-based repression in developmental gene regulation. We also leverage these data to link enhancers to putative target genes, revealing connections between coding and non-coding sequence variation in disease etiology. Our study provides a compendium of resources for biomedical researchers, and achieves the most comprehensive view of embryonic chromatin states to date.

genomics

The proBAM and proBed standard formats: enabling a seamless integration of genomics and proteomics data.

On behalf of The Human Proteome Organization (HUPO) Proteomics Standards Initiative (PSI), we are here introducing two novel standard data formats, proBAM and proBed, that have been developed to address the current challenges of integrating mass spectrometry based proteomics data with genomics and transcriptomics information in proteogenomics studies. proBAM and proBed are adaptations from the well-defined, widely used file formats SAM/BAM and BED respectively, and both have been extended to meet specific requirements entailed by proteomics data. Therefore, existing popular genomics tools such as SAMtools and Bedtools, and several very popular genome browsers, can be used to manipulate and visualize these formats already out-of-the-box. We also highlight that a number of specific additional software tools, properly supporting the proteomics information available in these formats, are now available providing functionalities such as file generation, file conversion, and data analysis. All the related documentation to the formats, including the detailed file format specifications, and example files are accessible at http://www.psidev.info/probam and http://www.psidev.info/probed.

bioinformatics

An Efficient Platform For Astrocyte Differentiation From Human Induced Pluripotent Stem Cells

Growing evidence implicates the importance of glia, particularly astrocytes, in neurological and psychiatric diseases. Here, we describe a rapid and robust method for the differentiation of highly pure populations of replicative astrocytes from human induced pluripotent stem cells (hiPSCs), via a neural progenitor cell (NPC) intermediate. Using this method, we generated hiPSC-derived astrocyte populations (hiPSC-astrocytes) from 42 NPC lines (derived from 30 individuals) with an average of [~]90% S100{beta}-positive cells. Transcriptomic analysis demonstrated that the hiPSC-astrocytes are highly similar to primary human fetal astrocytes and characteristic of a non-reactive state. hiPSC-astrocytes respond to inflammatory stimulants, display phagocytic capacity and enhance microglial phagocytosis. hiPSC-astrocytes also possess spontaneous calcium transient activity. Our novel protocol is a reproducible, straightforward (single media) and rapid (<30 days) method to generate homogenous populations of hiPSC-astrocytes that can be used for neuron-astrocyte and microglia-astrocyte co-cultures for the study of neuropsychiatric disorders.\n\nABBREVIATIONS

neuroscience

QuickRNASeq: Guide For Pipeline Implementation And For Interactive Results Visualization

i.Summary/AbstractSequencing of transcribed RNA molecules (RNA-seq) has been used wildly for studying cell transcriptomes in bulk or at the single-cell level (1, 2, 3) and is becoming the de facto technology for investigating gene expression level changes in various biological conditions, on the time course, and under drug treatments. Furthermore, RNA-Seq data helped identify fusion genes that are related to certain cancers (4). Differential gene expression before and after drug treatments provides insights to mechanism of action, pharmacodynamics of the drugs, and safety concerns (5). Because each RNA-seq run generates tens to hundreds of millions of short reads with size ranging from 50bp-200bp, a tool that deciphers these short reads to an integrated and digestible analysis report is in high demand. QuickRNASeq (6) is an application for large-scale RNA-seq data analysis and real-time interactive visualization of complex data sets. This application automates the use of several of the best open-source tools to efficiently generate user friendly, easy to share, and ready to publish report. Figure 1 illustrates some of the interactive plots produced by QuickRNASeq. The visualization features of the application have been further improved since its first publication in early 2016. The original QuickRNASeq publication (6) provided details of background, software selection, and implementation. Here, we outline the steps required to implement QuickRNASeq in users own environment, as well as demonstrate some basic yet powerful utilities of the advanced interactive visualization modules in the report.\n\nO_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC=\"FIGDIR/small/125856_fig1.gif\" ALT=\"Figure 1\">\nView larger version (59K):\norg.highwire.dtl.DTLVardef@1f6fb70org.highwire.dtl.DTLVardef@1f5a748org.highwire.dtl.DTLVardef@b990fborg.highwire.dtl.DTLVardef@dd5336_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 1C_FLOATNO Interactive plots from QuickRNASeq report. Figures (a, b, c) can be retrieved by clicking on the pointing hands as shown in figure Id. On any of these interactive plots, mouse over each sample displays associated sample QC metrics, (a) Read mapping summary in the expanded display mode, (b) SNP concordance matrix of 48 samples from 5 donors. Samples from the same donor should be highly concordant, (c) Gene expression chart, which shows the number of genes past various expression thresholds, (d) Center portion of the QuickRNASeq report, (e) Parallel plot linking multiple QC measures for the same samples plus table of multi-dimensional QC measures.\n\nC_FIG

bioinformatics