bioRxiv ScienceSearch

Biology subjects

Zhang, S.

Publications and source records attributed to Zhang, S..

85 records · Page 5Linked to original sources

Base-Specific Mutational Intolerance Near Splice-Sites Clarifies Role Of Non-Essential Splice Nucleotides

Variation in RNA splicing (i.e., alternative splicing) plays an important role in many diseases. Variants near 5' and 3' splice sites often affect splicing, but the effects of these variants on splicing and disease have not been fully characterized beyond the 2 \"essential\" splice nucleotides flanking each exon. Here we provide quantitative measurements of tolerance to mutational disruptions by position and reference allele-alternative allele combination. We show that certain reference alleles are particularly sensitive to mutations, regardless of the alternative alleles into which they are mutated. Using public RNA-seq data, we demonstrate that individuals carrying such variants have significantly lower levels of the correctly spliced transcript compared to individuals without them, and confirm that these specific substitutions are highly enriched for known Mendelian mutations. Our results propose a more refined definition of the \"splice region\" and offer a new way to prioritize and provide functional interpretation of variants identified in diagnostic sequencing and association studies.

genomics

Characterizing RNA Pseudouridylation By Convolutional Neural Networks

The most prevalent post-transcriptional RNA modification, pseudouridine ({Psi}), also known as the fifth ribonucleoside, is widespread in rRNAs, tRNAs, snRNAs, snoRNAs and mRNAs. Pseudouridines in RNAs are implicated in many aspects of post-transcriptional regulation, such as the maintenance of translation fidelity, control of RNA stability and stabilization of RNA structure. However, our understanding of the functions, mechanisms as well as precise distribution of pseudourdines (especially in mRNAs) still remains largely unclear. Though thousands of RNA pseudouridylation sites have been identified by high-throughput experimental techniques recently, the landscape of pseudouridines across the whole transcriptome has not yet been fully delineated. In this study, we present a highly effective model, called PULSE (PseudoUridyLation Sites Estimator), to predict novel {Psi} sites from large-scale profiling data of pseudouridines and characterize the contextual sequence features of pseudouridylation. PULSE employs a deep learning framework, called convolutional neural network (CNN), which has been successfully and widely used for sequence pattern discovery in the literature. Our extensive validation tests demonstrated that PULSE can outperform conventional learning models and achieve high prediction accuracy, thus enabling us to further characterize the transcriptome-wide landscape of pseudouridine sites. Overall, PULSE can provide a useful tool to further investigate the functional roles of pseudouridylation in post-transcriptional regulation.

bioinformatics

Varying Effects of Common Tuberculosis Drugs on Enhancing Clofazimine Activity in vitro

Clofazimine (CFZ), originally developed as an anti-tuberculosis (TB) drug in the 1950s 1, is commonly used to treat leprosy and also nontuberculous mycobacterial (NTM) infections. 2 Although CFZ has good activity against Mycobacterium tuberculosis, it was not used in the treatment of pulmonary TB mainly because it had the side effect of skin discoloration and there were other more effective drugs like isoniazid (INH), rifampin (RIF) and pyrazinamide (PZA) already available for the treatment of TB. 2 However, the increasing emergence of multi-drug-resistant TB (MDR-TB) has revived interest in the use of CFZ to treat MDR-TB. 2,3\n\nAlthough resistance to CFZ has been shown to be mediated by mutations in Rv0678,4,5 Rv1979c, or Rv2535c (PepQ),5 the mode of action of CFZ has remained poorly understood. CFZ appears to hav ...

microbiology

Identification of novel mutations associated with cycloserine resistance in Mycobacterium tuberculosis

ObjectivesD-cycloserine (DCS) is an important second-line drug used to treat multi-drug resistant (MDR) and extensively drug-resistant (XDR) tuberculosis. However, the mechanisms of resistance to DCS are not well understood. Here we investigated the molecular basis of DCS resistance using in vitro isolated resistant mutants of Mycobacterium tuberculosis.\n\nMethodsM. tuberculosis H37Rv was subjected to mutant selection on 7H11 agar plates containing varying concentrations of DCS. A total of 35 DCS-resistant mutants were isolated and 18 mutants were subjected to whole genome sequencing. The identified mutations associated with DCS resistance were confirmed by PCR-Sanger sequencing.\n\nResultsWe identified mutations in 17 genes that are associated with DCS resistance. Except mutations in alr (rv3423c) which is known to be involved in DCS resistance, 16 new genes rv0059, betP (rv0917), rv0221, rv1403c, rv1683, rv1726, gabD2 (rv1731), rv2749, sugI (rv3331), hisC2 (rv3 772), single mutation in 5 intergenic region of rv3345c and rv1435c, and insertion in 3 region of rv0759c were identified as solo mutations in their respective DCS-resistant mutants. Our findings indicate that the mechanisms of DCS resistance are more complex than previously thought and involve genes participating in different cellular functions such as lipid metabolism, methyltransferase, stress response, and transport proteins.\n\nConclusionsNew mutations in diverse genes associated with DCS are identified, which shed new light on the mechanisms of action and resistance of DCS. Future studies are needed to verify these findings in clinical strains so that molecular detection of DCS resistance for improved treatment of MDR-TB can be developed.

microbiology

Identification of drug candidates that enhance pyrazinamide activity from a clinical drug library

Tuberculosis (TB) remains a leading cause of morbidity and mortality globally despite the availability of the TB therapy. 1 The current TB therapy is lengthy and suboptimal, requiring a treatment time of at least 6 months for drug susceptible TB and 9-12 months (shorter Bangladesh regimen) or 18-24 months (regular regimen) for multi-drug-resistant tuberculosis (MDR-TB). 1 The lengthy therapy makes patient compliance difficult, which frequently leads to emergence of drug-resistant strains. The requirement for the prolonged treatment is thought to be due to dormant persister bacteria which are not effectively killed by the current TB drugs, except rifampin and pyrazinamide (PZA) which have higher activity against persisters. 2, 3 Therefore new therapies should address the problem of insufficient efficacy against M. tuberculosis persisters, which could cause relapse of clinical disease. 4 PZA is a critical frontline TB drug that kills persister bacteria 5 and shortens the TB treatment from 9-12 months to 6 months. 6, 7 Although several new TB drugs are showing promise in clinical studies, none can replace PZA as they all have to be used together with PZA. 7 Because of the essentiality of PZA and the high cost of developing new drugs, in this study, we explored the idea of identifying drugs that enhance the anti-persister activity of PZA as an economic alternative approach to developing new drugs for improved treatment by screening an clinical drug library against old M. tuberculosis cultures enriched with persisters.

microbiology

Decoding temporal interpretation of the morphogen Bicoid in the early Drosophila embryo

Morphogen gradients provide essential spatial information during development. Not only the local concentration but also duration of morphogen exposure is critical for correct cell fate decisions. Yet, how and when cells temporally integrate signals from a morphogen remains unclear. Here, we use optogenetic manipulation to switch off Bicoid-dependent transcription in the early Drosophila embryo with high temporal resolution, allowing time-specific and reversible manipulation of morphogen signalling. We find that Bicoid transcriptional activity is dispensable for embryonic viability in the first hour after fertilization, but persistently required throughout the rest of the blastoderm stage. Short interruptions of Bicoid activity alter the most anterior cell fate decisions, while prolonged inactivation expands patterning defects from anterior to posterior. Such anterior susceptibility correlates with high reliance of anterior gap gene expression on Bicoid. Therefore, cell fates exposed to higher Bicoid concentration require input for longer duration, demonstrating a previously unknown aspect of morphogen decoding.

developmental biology

Zebrafish models for human ALA-dehydratase-deficient porphyria (ADP) and hereditary coproporphyria (HCP) generated with TALEN and CRISPR-Cas9

Defects in the enzymes involved in heme biosynthesis result in a group of human metabolic genetic disorders known as porphyrias. Using a zebrafish model for human hepatoerythropoietic porphyria (HEP), caused by defective uroporphyrinogen decarboxylase (Urod), the fifth enzyme in the heme biosynthesis pathway, we recently have found a novel aspect of porphyria pathogenesis. However, no hereditable zebrafish models with genetic mutations of alad and cpox, encoding the second enzyme delta-aminolevulinate dehydratase (Alad) and the sixth enzyme coproporphyrinogen oxidase (Cpox), have been established to date. Here we employed site-specific genome-editing tools transcription activator-like effector nuclease (TALEN) and clustered regularly interspaced short palindromic repeats (CRISPR)/CRISPR-associated protein 9 (Cas9) to generate zebrafish mutants for alad and cpox. These zebrafish mutants display phenotypes of heme deficiency, hypochromia, abnormal erythrocytic maturation and accumulation of heme precursor intermediates, reminiscent of human ALA-dehydratase-deficient porphyria (ADP) and hereditary coproporphyrian (HCP), respectively. Further, we observed altered expression of genes involved in heme biosynthesis and degradation and particularly down-regulation of exocrine pancreatic zymogens in ADP (alad-/-) and HCP (cpox-/-) fishes. These two zebrafish porphyria models can survive at least 7 days and thus provide invaluable resources for elucidating novel pathological aspects of porphyrias, evaluating mutated forms of human ALAD and CPOX, discovering new therapeutic targets and developing effective drugs for these complex genetic diseases. Our studies also highlight generation of zebrafish models for human diseases with two versatile genome-editing tools.

genetics

Ultra-fast Identity by Descent Detection in Biobank-Scale Cohorts using Positional Burrows-Wheeler Transform

With the availability of genotyping data of very large samples, there is an increasing need for tools that can efficiently identify genetic relationships among all individuals in the sample. One fundamental measure of genetic relationship of a pair of individuals is identity by descent (IBD), chromosomal segments that are shared among two individuals due to common ancestry. However, the efficient identification of IBD segments among a large number of genotyped individuals is a challenging computational problem. Most existing methods are not feasible for even thousands of individuals because they are based on pairwise comparisons of all individuals and thus scale up quadratically with sample size. Some methods, such as GERMLINE, use fast dictionary lookup of short seed sequence matches to achieve a near-linear time efficiency. However, the number of short seed matches often scales up super-linearly in real population data.\n\nIn this paper we describe a new approach for IBD detection. We take advantage of an efficient population genotype index, Positional BWT (PBWT), by Richard Durbin. PBWT achieves linear time query of perfectly identical subsequences among all samples. However, the original PBWT is not tolerant to genotyping errors which often interrupt long IBD segments into short fragments. We introduce a randomized strategy by running PBWTs over random projections of the original sequences. To boost the detection power we run PBWT multiple times and merge the identified IBD segments through interval tree algorithms. Given a target IBD segment length, RaPID adjust parameters to optimize detection power and accuracy.\n\nSimulation results proved that our tool (RaPID) achieves almost linear scaling up to sample size and is orders of magnitude faster than GERMLINE. At the same time, RaPID maintains a detection power and accuracy comparable to existing mainstream algorithms, GERMLINE and IBDseq. Running multiple times with various target detection lengths over the 1000 Genomes Project data, RaPID can detect population events at different time scales. With our tool, it is feasible to identify IBDs among hundreds of thousands to millions of individuals, a sample size that will become reality in a few years.

genomics

TIDE: predicting translation initiation sites by deep learning

MotivationTranslation initiation is a key step in the regulation of gene expression. In addition to the annotated translation initiation sites (TISs), the translation process may also start at multiple alternative TISs (including both AUG and non-AUG codons), which makes it challenging to predict TISs and study the underlying regulatory mechanisms. Meanwhile, the advent of several high-throughput sequencing techniques for profiling initiating ribosomes at single-nucleotide resolution, e.g., GTI-seq and QTI-seq, provides abundant data for systematically studying the general principles of translation initiation and the development of computational method for TIS identification.\n\nMethodsWe have developed a deep learning based framework, named TITER, for accurately predicting TISs on a genome-wide scale based on QTI-seq data. TITER extracts the sequence features of translation initiation from the surrounding sequence contexts of TISs using a hybrid neural network and further integrates the prior preference of TIS codon composition into a unified prediction framework.\n\nResultsExtensive tests demonstrated that TITER can greatly outperform the state-of-the-art prediction methods in identifying TISs. In addition, TITER was able to identify important sequence signatures for individual types of TIS codons, including a Kozak-sequence-like motif for AUG start codon. Furthermore, the TITER prediction score can be related to the strength of translation initiation in various biological scenarios, including the repressive effect of the upstream open reading frames (uORFs) on gene expression and the mutational effects influencing translation initiation efficiency.\n\nAvailabilityTITER is available as an open-source software and can be downloaded from https://github.com/zhangsaithu/titer\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

Deficiency of Voltage-gated Proton Channel Hv1 Leads to hypoinsulinaemia, hyperglycemia and glucose intolerance in mice

Here, we demonstrate that the voltage-gated proton channel Hv1 represents a regulatory mechanism for insulin secretion of pancreatic islet {beta} cell. In vivo, Hv1-de[fi]cient mice display hyperglycemia and glucose intolerance due to reduced insulin secretion, but normal peripheral insulin sensitivity. In vitro, islets of Hv1-de[fi]cient and heterozygous mice, INS-1 (832/13) cells with siRNA-mediated knockdown of Hv1 exhibit a marked defect in glucose- and K+-induced insulin secretion. Hv1 de[fi]ciency decreases both insulin and proinsulin contents, and limits glucose-induced Ca2+ entry and membrane depolarization. Furthermore, loss of Hv1 increases insulin-containing granular pH and decreases cytosolic pH. In addition, histologic studies show a decrease in {beta} cell mass in islets of Hv1-deficient mice. Collectively, our results indicate that Hv1 supports insulin secretion in the {beta} cell by calcium entry, membrane depolarization and intracellular pH regulation.

cell biology

Heat-stable preservation of protein expression systems for portable therapeutics production

Many biotechnology capabilities are limited by stringent storage needs of reagents, largely prohibiting use outside of specialized laboratories. Focusing on a large class of protein-based biotechnology applications, we address this issue by developing a method for preserving cell-free protein expression systems under months of heat stress. Our approach realizes an unprecedented degree of long term heat stability by leveraging the sugar alcohol trehalose, a simple, low-cost, open-air drying step, and strategic separation of sets of reaction components during drying. The resulting preservation capacity opens the door for efficient production of a wide range of on-demand proteins under adverse conditions, for instance during emergency outbreaks or in remote or otherwise inaccessible locations. As such, our preservation method stands to advance a great number of different cell-free technologies, including remediation efforts, point of care therapeutics, and large-scale biosensing. To demonstrate this application potential, we use cell-free reagents subjected to months of heat stress and atmospheric conditions to produce sufficient concentrations of a pyocin protein to kill Pseudomonas aeruginosa, one of the most troublesome pathogens for traumatic and burn wound injuries. Our work makes possible new biotechnology applications that demand both ruggedness and scalability.

synthetic biology

A Deep Boosting Based Approach for Capturing the Sequence Binding Preferences of RNA-Binding Proteins from High-Throughput CLIP-Seq Data

Characterizing the binding behaviors of RNA-binding proteins (RBPs) is important for understanding their functional roles in gene expression regulation. However, current high-throughput experimental methods for identifying RBP targets, such as CLIP-seq and RNAcompete, usually suffer from the false positive and false negative issues. Here, we develop a deep boosting based machine learning approach, called DeBooster, to accurately model the binding sequence preferences and identify the corresponding binding targets of RBPs from CLIP-seq data. Comprehensive validation tests have shown that DeBooster can outperform other state-of-the-art approaches in predicting RBP targets and recover false negatives that are common in current CLIP-seq data. In addition, we have demonstrated several new potential applications of DeBooster in understanding the regulatory functions of RBPs, including the binding effects of the RNA helicase MOV10 on mRNA degradation, the influence of different binding behaviors of the ADAR proteins on RNA editing, as well as the antagonizing effect of RBP binding on miRNA repression. Moreover, DeBooster may provide an effective index to investigate the effect of pathogenic mutations in RBP binding sites, especially those related to splicing events. We expect that DeBooster will be widely applied to analyze large-scale CLIP-seq experimental data and can provide a practically useful tool for novel biological discoveries in understanding the regulatory mechanisms of RBPs. The scource code of DeBooster can be downloaded from http://github.com/dongfanghong/deepboost.

bioinformatics

Chromosomal dynamics predicted by an elastic network model explains genome-wide accessibility and long-range couplings

Understanding the three-dimensional (3D) architecture of the chromatin and its relation to gene expression and regulation is fundamental to understanding how the genome functions. Advances in Hi-C technology now permit us to have a glimpse into the 3D genome organization and identify topologically associated domains (TADs), but we still lack an understanding of the structural dynamics of chromosomes. The dynamic couplings between regions separated by large genomic distances (> 50 megabases) have yet to be characterized. We adapted a well-established protein-modeling framework, the Gaussian Network Model (GNM), to the task of modeling chromatin dynamics using Hi-C contact data. We show that the GNM can identify structural dynamics at multiple scales: it can quantify the fluctuations in the positions of gene loci, find large genomic compartments and smaller TADs that undergo en-bloc movements, and identify dynamically coupled distal regions along the chromosomes. We show that the predictions of the GNM correlate well with DNase-seq and ATAC-seq measurements on accessibility, the previously identified A and B compartments of chromatin structure, and pairs of interacting loci identified by ChIA-PET. We describe a method to use the GNM to identify novel cross-correlated distal domains (CCDDs) representing regions of long-range dynamic coupling and show that CCDDs are often associated with increased gene coexpression using a large-scale analysis of 212 expression experiments. Together, these results show that GNM provides a mathematically well-founded unified framework for assessing chromatin dynamics and the structural basis of genome-wide observations.

genomics