bioRxiv ScienceSearch

Biology subjects

Zhang, L.

Publications and source records attributed to Zhang, L..

At least 73 records · Page 4Linked to original sources

GDCRNATools: an R/Bioconductor package for integrative analysis of lncRNA, miRNA, and mRNA data in GDC

The large-scale multidimensional omics data in the Genomic Data Commons (GDC) provides opportunities to investigate the crosstalk among different RNA species and their regulatory mechanisms in cancers. Easy-to-use bioinformatics pipelines are needed to facilitate such studies. We have developed a user-friendly R/Bioconductor package, named GDCRNATools, to facilitate downloading, organizing, and analyzing RNA data in GDC with an emphasis on deciphering the lncRNA-mRNA related competing endogenous RNAs (ceRNAs) regulatory network in cancers. Many widely used bioinformatics tools and databases are utilized in our package. Users can easily pack preferred downstream analysis pipelines or integrate their own pipelines into the workflow. Interactive shiny web apps built in GDCRNATools greatly improve visualization of results from the analysis.\n\nAvailabilityGDCRNATools is an R/Bioconductor package that is freely available at https://github.com/Jialab-UCR/GDCRNATools

bioinformatics

Highly tamoxifen-inducible principal-cell-specific Cre mice with complete fidelity in cell specificity and no leakiness

An ideal inducible system should be cell-specific and have absolute no background recombination without induction (i.e. no leakiness), a high recombination rate after induction, and complete fidelity in cell specificity (i.e. restricted recombination exclusively in cells where the driver gene is expressed). However, such an ideal mouse model remains unavailable for collecting duct research. Here, we report a mouse model that meets these criteria. In this model, a cassette expressing ERT2CreERT2 (ECE) is inserted at the ATG of the endogenous Aqp2 locus to disrupt Aqp2 function and to express ECE under the control of the Aqp2 promoter. The resulting allele is named Aqp2ECE. There was no indication of a significant impact of disruption of a copy of Aqp2 on renal function and blood pressure control in adult Aqp2ECE/+ heterozygotes. Without tamoxifen, Aqp2ECE did not activate a Cre-dependent red fluorescence protein (RFP) reporter in adult kidneys. A single injection of tamoxifen (2 mg) to adult mice enables Aqp2ECE to induce robust RFP expression in the whole kidney 24h post injection, with the highest recombination efficiency of 95% in the inner medulla. All RFP-labeled cells expressed principal cell markers (Aqp2 & Aqp3), but not intercalated cell markers (V-ATPase B1B2, and carbonic anhydrase II). Hence, Aqp2ECE confers principal cell-specific tamoxifen-inducible recombination with absolute no leakiness, high inducibility, and complete fidelity in cell specificity, which should be an important tool for temporospatial control of target genes in the principal cells and for Aqp2+ lineage tracing in adult mice.

genetics

CaDrA: A computational framework for performing candidate driver analyses using binary genomic features

Identifying complementary genetic drivers of a given phenotypic outcome is a challenging task that is important to gaining new biological insight and discovering targets for disease therapy. Existing methods aimed at achieving this task lack analytical flexibility. We developed Candidate Driver Analysis or CaDrA, a framework to identify functionally-relevant subsets of binary genomic features that, together, are associated with a specific outcome of interest. We evaluate CaDrAs sensitivity and specificity for typically-sized multi-omic datasets, and demonstrate CaDrAs ability to identify both known and novel drivers of oncogenic activity in cancer cell lines and primary tumors.

bioinformatics

HAPDeNovo: a haplotype-based approach for filtering and phasing de novo mutations in linked read sequencing data

BackgroundDe novo mutations (DNMs) are associated with neurodevelopmental and congenital diseases, and their detection can contribute to understanding disease pathogenicity. However, accurate detection is challenging because of their small number relative to the genome-wide false positives in next generation sequencing (NGS) data. Software such as DeNovoGear and TrioDeNovo have been developed to detect DNMs, but at good sensitivity they still produce many false positive calls.\n\nResultsTo address this challenge, we develop HAPDeNovo, a program that leverages phasing information from linked read sequencing, to remove false positive DNMs from candidate lists generated by DNM-detection tools. Short reads from each phasing block are allocated to each of the two haplotypes followed by generating a haploid genotype for each putative DNM.HAPDeNovo removes variants that are called as heterozygous in one of the haplotypes because they are almost certainly false positives. Our experiments on 10X Chromium linked read sequencing trio data reveal that HAPDeNovo eliminates 80% to 99% of false positives regardless of how large the candidate DNM set is.\n\nConclusionsHAPDeNovo leverages the haplotype information from linked read sequencing to remove spurious false positive DNMs effectively, and it increases accuracy of DNM detection dramatically without sacrificing sensitivity.

genomics

YlmD and YlmE are required for correct sporulation-specific cell division in Streptomyces coelicolor A3(2)

Cell division during the reproductive phase of the Streptomyces life-cycle requires tight coordination between synchronous formation of multiple septa and DNA segregation. One remarkable difference with most other bacterial systems is that cell division in Streptomyces is positively controlled by the recruitment of FtsZ by SsgB. Here we show that deletion of ylmD (SCO2081) or ylmE (SCO2080), which lie in operon with ftsZ in the dcw cluster of actinomycetes, has major consequences for sporulation-specific cell division in Streptomyces coelicolor. Electron and fluorescence microscopy demonstrated that ylmE mutants have a highly aberrant phenotype with defective septum synthesis, and produce very few spores with low viability and high heat sensitivity. FtsZ-ring formation was also highly disturbed in ylmE mutants. Deletion of ylmD had a far less severe effect on sporulation. Interestingly, the additional deletion of ylmD restored sporulation to the ylmE null mutant. YlmD and YlmE are not part of the divisome, but instead localize diffusely in aerial hyphae, with differential intensity throughout the sporogenic part of the hyphae. Taken together, our work reveals a function for YlmD and YlmE in the control of sporulation-specific cell division in S. coelicolor, whereby the presence of YlmD alone results in major developmental defects.

microbiology

Smad9 is a key player of follicular selection in goose via keeping the balance of LHR transcription

The egg production of poultry depends on follicular development and selection. However, the mechanism of selecting the priority of hierarchical follicles is completely unknown. Smad9 is one of the important transcription factors in BMP/Smads pathway and involved in goose follicular initiation. To explore its potential role in goose follicle hierarchy determination, we first blocked Smad9 expression using BMP typereceptor inhibitor LDN-193189 both in vivo and in vitro. Unexpectedly, LDN-193189 administration could dramatically suppress Smad9 level and elevate egg production (7.08 eggs / bird, P< 0.05) of animals, and the estradiol (E2) and luteinizing hormone receptor (LHR) level were significantly increased (P< 0.05), but the progesterone (P4) and follicle stimulating hormone receptor (FSHR) mRNA remain unchanged. Surprisingly, Smad9 knockdown notably attenuated (P< 0.05) in E2, P4, FSHR and LHR level in goose granulosa cells (gGCs). Further chromatin immunoprecipitation (ChIP) assay of gGCs revealed that Smad9, served as a sensor of balance, bound to the LHR promoter regulating its transcription. These findings demonstrated that Smad9 is differentially expressed in goose follicles, and acts as a key player in controlling goose follicular selection.\n\nSUMMARY STATEMENTTo study the hierarchical development mechanism of avian follicle, new strategies can be found to improve the egg production of low-yielding poultry, such as geese.

developmental biology

The DEAD box RNA helicase Ddx39a is essential for myocyte and lens development in zebrafish

RNA helicases from the DEAD-box family are found in almost all organisms and have important roles in RNA metabolism including RNA synthesis, processing and degradation. The function and mechanism of action of most of these helicases in animal development and human disease are largely unexplored. In a zebrafish mutagenesis screen to identify genes essential for heart development we identified a zebrafish mutant, which disrupts the gene encoding the RNA helicase DEAD-box 39a (ddx39a).Homozygous ddx39a mutant embryos exhibit profound cardiac and trunk muscle dystrophy, along with lens abnormalities caused by abrupt terminal differentiation of cardiomyocyte, myoblast and lens fiber cells. Further investigation indicated that loss of ddx39a hindered mRNA splicing of members of the kmt2 gene family, leading to mis-regulation of structural gene expression in cardiomyocyte, myoblast and lens fiber cells. Taken together, these results show that Ddx39a plays an essential role in establishment of proper epigenetic status during cell differentiation.

developmental biology

InfoTrim: A DNA Read Quality Trimmer Using Entropy

Biological DNA reads are often trimmed before mapping, genome assembly, and other tasks to improve the quality of the results. Biological sequence complexity relates to alignment quality as low complexity regions can align poorly. There are many read trimmers, but many do not use sequence complexity for trimming. Alignment of reads generated from whole genome bisulfite sequencing is especially challenging since bisulfite treated reads tend to reduce sequence complexity. InfoTrim, a new read trimmer, was created to explore these issues. It is evaluated against five other trimmers using four read mappers on real and simulated bisulfite treated DNA data. InfoTrim produces reasonable results consistent with other trimmers.

bioinformatics

Biopipe: A Lightweight System Enabling Comparison of Bioinformatics Tools and Workflows

Analyzing next generation sequencing data always requires researchers to install many tools, prepare input data compliant to the required data format, and execute the tools in specific orders. Such tool installation and workflow execution process is tedious and error-prone, and becomes very challenging when researchers need to compare multiple alternative tool chains. To mitigate this problem, we developed a new lightweight and portable system, Biopipe, to simplify the creation and execution of bioinformatics tools and workflows, and to further enable the comparison between alternative tools or workflows. Biopipe allows users to create and edit workflows with user-friendly web interfaces, and automates tool installation as well as workflow synthesis by downloading and executing predefined Docker images. With Biopipe, biologists can easily experiment with and compare different bioinformatics tools and workflows without much computer science knowledge. There are mainly two parts in Biopipe: a web application and a standalone Java application. They are freely available at http://bench.cs.vt.edu:8282/Biopipe-Workflow-Editor-0.0.1/index.xhtml and https://code.vt.edu/saima5/Biopipe-Run-Workflow\n\nContactnm8247@cs.vt.edu\n\nSupplementary informationSupplementary data are available online.

bioinformatics

FastViromeExplorer: A Pipeline for Virus and Phage Identification and Abundance Profiling in Metagenomics Data

Identifying viruses and phages in a metagenomics sample has important implication in improving human health, preventing viral outbreaks, and developing personalized medicine. With the rapid increase in data files generated by next generation sequencing, existing tools for identifying and annotating viruses and phages in metagenomics samples suffer from expensive running time. In this paper, we developed a stand-alone pipeline, FastViromeExplorer, for rapid identification and abundance quantification of viruses and phages in big metagenomic data. Both real and simulated data validated FastViromeExplorer as a reliable tool to accurately identify viruses and their abundances in large data, as well as in a time efficient manner.

bioinformatics

Allosteric effector ppGpp potentiates the inhibition of transcript initiation by DksA

DksA and ppGpp are the central players in the Escherichia coli stringent response and mediate a complete reprogramming of the transcriptome from one optimized for rapid growth to one adapted for survival during nutrient limitation. A major component of the response is a reduction in ribosome synthesis, which is accomplished by the synergistic action of DksA and ppGpp bound to RNA polymerase (RNAP) inhibiting transcription of rRNAs. Here, we report the X-ray crystal structures of E. coli RNAP holoenzyme in complex with DksA alone and with ppGpp. The structures show that DksA accesses the template strand at the active site and the downstream DNA binding site of RNAP simultaneously and reveal that binding of the allosteric effector ppGpp reshapes the RNAP-DksA complex. The structural data support a model for transcriptional inhibition in which ppGpp potentiates the destabilization of open complexes on rRNA promoters by DksA. We also determined the structure of RNAP-TraR complex, which reveals the mechanism of ppGpp-independent transcription inhibition by TraR. This work establishes new ground for understanding the pleiotropic effects of DksA and ppGpp on transcriptional regulation in proteobacteria.\n\nHighlightsO_LIDksA has two modes of binding to RNA polymerase\nC_LIO_LIDksA is capable of inhibiting the catalysis and influences the DNA binding of RNAP\nC_LIO_LIppGpp acts as an allosteric effector of DksA function\nC_LIO_LIppGpp stabilizes DksA in a more functionally important binding mode\nC_LI

molecular biology

SelexGLM differentiates androgen and glucocorticoid receptor DNA-binding preference over an extended binding site

The DNA-binding interfaces of the androgen (AR) and glucocorticoid (GR) receptors are virtually identical, yet these transcription factors share only about a third of their genomic binding sites and regulate similarly distinct sets of target genes. To address this paradox, we determined the intrinsic specificities of the AR and GR DNA binding domains using a refined version of SELEX-seq. We developed an algorithm, SelexGLM, that quantifies binding specificity over a large (31 bp) binding-site by iteratively fitting a feature-based generalized linear model to SELEX probe counts. This analysis revealed that the DNA binding preferences of AR and GR homodimers differ significantly, both within and outside the 15bp core binding site. The relative preference between the two factors can be tuned over a wide range by changing the DNA sequence, with AR more sensitive to sequence changes than GR. The specificity of AR extends to the regions flanking the core 15bp site, where isothermal calorimetry measurements reveal that affinity is augmented by enthalpy-driven readout of poly-A sequences associated with narrowed minor groove width. We conclude that the increased specificity of AR is correlated with more enthalpy-driven binding than GR. The binding models help explain differences in AR and GR genomic binding, and provide a biophysical rationale for how promiscuous binding by GR allows functional substitution for AR in some castration-resistant prostate cancers.

systems biology

The effect of sequence mismatches on binding affinity and endonuclease activity are decoupled throughout the Cas9 binding site

The CRISPR-Cas9 system is a powerful genomic tool. Although targeted to complementary genomic sequences by a guide RNA (gRNA), Cas9 tolerates gRNA:DNA mismatches and cleaves off-target sites. How mismatches quantitatively affect Cas9 binding and cutting is not understood. Using SelexGLM to construct a comprehensive model for DNA-binding specificity, we observed that 13-bp of complementarity in the PAM-proximal DNA contributes to affinity. We then adapted Spec-seq and developed SEAM-seq to systematically compare the impact of gRNA:DNA mismatches on affinity and endonuclease activity, respectively. Though most often coupled, these simple and accessible experiments identified sometimes opposing effects for mismatches on DNA-binding and cutting. In the PAM-distal region mismatches decreased activity but not affinity, whereas in the PAM-proximal region some reduced-affinity mismatches enhanced activity. This mismatch-activation was particularly evident where the gRNA:DNA duplex bends. We developed integrative models from these measurements that estimate catalytic efficiency and can be used to predict off-target cleavage.

biochemistry

Uncovering Medical Insights from Vast Amounts of Biomedical Data in Clinical Case Reports

Clinical case reports (CCRs) have a time-honored tradition in serving as an important means of sharing clinical experiences on patients presenting with atypical disease phenotypes or receiving new therapies. However, the huge amount of accumulated case reports are isolated, unstructured, and heterogeneous clinical data, posing a great challenge to clinicians and researchers in mining relevant information through existing indexing tools. In this investigation, in order to render CCRs more findable, accessible, interoperable, and reusable (FAIR) by the biomedical community, we created a resource platform, including the construction of a test dataset consisting of 1000 CCRs spanning 14 disease phenotypes, a standardized metadata template and metrics, and a set of computational tools to automatically retrieve relevant medical information and to analyze all published PubMed clinical case reports with respect to trends in publication journals, citations impact, MeSH Terms, drug use, distributions of patient demographics, and relationships with other case reports and databases. Our standardized metadata template and CCR test dataset may be valuable resources to advance medical science and improve patient care for researchers who are using machine learning approaches with a high-quality dataset to train and validate their algorithms. In the future, our analytical tools may be applied towards other large clinical data sources as well.

bioinformatics

A compartmentalized, self-extinguishing signaling network mediates crossover control in meiosis

Meiotic recombination between homologous chromosomes is tightly regulated to ensure proper chromosome segregation. Each chromosome pair typically undergoes at least one crossover event (crossover assurance) but these exchanges are also strictly limited in number and widely spaced along chromosomes (crossover interference). This has implied the existence of chromosome-wide signals that regulate crossovers, but their molecular basis remains mysterious. Here we characterize a family of four related RING finger proteins in C. elegans. These proteins are recruited to the synaptonemal complex between paired homologs, where they act as two heterodimeric complexes, likely as E3 ubiquitin ligases. Genetic and cytological analysis reveals that they act with additional components to create a self-extinguishing circuit that controls crossover designation and maturation. These proteins also act at the top of a hierarchical chromosome remodeling process that enables crossovers to direct stepwise segregation. Work in diverse phyla indicates that related mechanisms mediate crossover control across eukaryotes.

cell biology

CRISPR/Cas9-mediated Knock-in of an Optimized TetO Repeat for Live Cell Imaging of Endogenous Loci

Nuclear organization has an important role in determining genome function; however, it is not clear how spatiotemporal organization of the genome relates to functionality. To elucidate this relationship, a high-throughput method for tracking any locus of interest is desirable. Here, we report an efficient and scalable method named SHACKTeR (Short Homology and CRISPR/Cas9-mediated Knock-in of a TetO Repeat) for live cell imaging of specific chromosomal regions. Compared to alternatives, our method does not require a nearby repetitive sequence and it requires only two modifications to the genome: CRISPR/Cas9-mediated knock-in of an optimized TetO repeat and its visualization by TetR-EGFP expression. Our simplified knock-in protocol, utilizing short homology arms integrated by PCR, was successful at labeling 9 different loci in HCT116 cells with up to 20% efficiency. These loci included both nuclear speckle-associated, euchromatin regions and nuclear lamina-associated, heterochromatin regions. We anticipate the general applicability and scalability of our method will enhance causative analyses between gene function and compartmentalization in a high-throughput manner.

cell biology

Crosslinkers both drive and brake cytoskeletal remodeling and furrowing in cytokinesis

Cytokinesis and other cell shape changes are driven by the actomyosin contractile cytoskeleton. The molecular rearrangements that bring about contractility in non-muscle cells are currently debated. Specifically, both filament sliding by myosin motors, as well as cytoskeletal crosslinking by myosins and non-motor crosslinkers, are thought to promote contractility. Here, we examined how the abundance of motor and non-motor crosslinkers controls the speed of cytokinetic furrowing. We built a minimal model to simulate the contractile dynamics of the C. elegans zygote cytokinetic ring. This model predicted that intermediate levels of non-motor crosslinkers would allow maximal contraction speed, which we found to be the case for the scaffold protein anillin, in vivo. Our model also demonstrated a non-linear relationship between the abundance of motor ensembles and contraction speed. In vivo, thorough depletion of non-muscle myosin II delayed furrow initiation, slowed F-actin alignment, and reduced maximum contraction speed, but partial depletion allowed faster-than-expected kinetics. Thus, both motor and non-motor crosslinkers promote cytokinetic ring closure when present at low levels, but act as a brake when present at higher levels. Together, our findings extend the growing appreciation for the roles of crosslinkers, but reveal that they not only drive but also brake cytoskeletal remodeling.

cell biology

DeepARG: A deep learning approach for predicting antibiotic resistance genes from metagenomic data

Growing concerns regarding increasing rates of antibiotic resistance call for global monitoring efforts. Monitoring of environmental media (e.g., wastewater, agricultural waste, food, and water) is of particular interest as these media can serve as sources of potential novel antibiotic resistance genes (ARGs), as hot spots for ARG exchange, and as pathways for the spread of ARGs and human exposure. Next-generation sequence-based monitoring has recently enabled direct access and profiling of the total metagenomic DNA pool, where ARGs are identified or predicted based on the \"best hits\" of homology searches against existing databases. Unfortunately, this approach tends to produce high rates of false negatives. To address such limitations, we propose here a deep leaning approach, taking into account a dissimilarity matrix created using all known categories of ARGs. Two models, deepARG-SS and deepARG-LS, were constructed for short read sequences and full gene length sequences, respectively. Performance evaluation of the deep learning models over 30 classes of antibiotics demonstrates that the deepARG models can predict ARGs with both high precision (>0.97) and recall (>0.90) for most of the antibiotic resistance categories. The models show advantage over the traditional best hit approach by having consistently much lower false negative rates and thus higher overall recall (>0.9). As more data become available for under-represented antibiotic resistance categories, the deepARG models performance can be expected to be further enhanced due to the nature of the underlying neural networks. The deepARG models are available both in command line version and via a Web server at http://bench.cs.vt.edu/deeparg. Our newly developed ARG database, deepARG-DB, containing predicted ARGs with high confidence and high degree of manual curation, greatly expands the current ARG repository. DeepARG-DB can be downloaded freely to benefit community research and future development of antibiotic resistance-related resources.\n\nAbbreviations

bioinformatics