bioRxiv ScienceSearch

Biology subjects

Uyar, B.

Publications and source records attributed to Uyar, B..

4 recordsLinked to original sources

Reproducible genomics analysis pipelines with GNU Guix

In bioinformatics, as well as other computationally-intensive research fields, there is a need for workflows that can reliably produce consistent output, independent of the software environment or configuration settings of the machine on which they are executed. Indeed, this is essential for controlled comparison between different observations or for the wider dissemination of workflows. Providing this type of reproducibility, however, is often complicated by the need to accommodate the myriad dependencies included in a larger body of software, each of which generally come in various versions. Moreover, in many fields (bioinformatics being a prime example), these versions are subject to continual change due to rapidly evolving technologies, further complicating problems related to reproducibility. Here, we propose a principled approach for building analysis pipelines and managing their dependencies. As a case study to demonstrate the utility of our approach, we present a set of highly reproducible pipelines for the analysis of RNA-seq, ChIP-seq, Bisulfite-seq, and single-cell RNA-seq. All pipelines process raw experimental data, and generate reports containing publication-ready plots and figures, with interactive report elements and standard observables. Users may install these highly reproducible packages and apply them to their own datasets without any special computational expertise beyond the use of the command line. We hope such a toolkit will provide immediate benefit to laboratory workers wishing to process their own data sets or bioinformaticians seeking to automate all, or parts of, their analyses. In the long term, we hope our approach to reproducibility will serve as a blueprint for reproducible workflows in other areas. Our pipelines, along with their corresponding documentation and sample reports, are available at http://bioinformatics.mdc-berlin.de/pigx

bioinformatics

FACT sets a barrier for cell fate reprogramming in C. elegans and Human

O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC=\"FIGDIR/small/185116_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (59K):\norg.highwire.dtl.DTLVardef@15ae7aaorg.highwire.dtl.DTLVardef@11f5c08org.highwire.dtl.DTLVardef@1d32442org.highwire.dtl.DTLVardef@f1808f_HPS_FORMAT_FIGEXP M_FIG C_FIG The chromatin regulator FACT (Facilitates Chromatin Transcription) is essential for ensuring stable gene expression by promoting transcription. In a genetic screen using C. elegans we identified that FACT maintains cell identities and acts as a barrier for transcription factor-mediated cell fate reprogramming. Strikingly, FACTs role as a reprogramming barrier is conserved in humans as we show that FACT depletion enhances reprogramming of fibroblasts into stem cells and neurons. Such activity of FACT is unexpected since known reprogramming barriers typically repress gene expression by silencing chromatin. In contrast, FACT is a positive regulator of gene expression suggesting an unprecedented link of cell fate maintenance with counteracting alternative cell identities. This notion is supported by ATAC-seq analysis showing that FACT depletion results in decreased but also increased chromatin accessibility for transcription factors. Our findings identify FACT as a cellular reprogramming barrier in C. elegans and humans, revealing an evolutionarily conserved mechanism for cell fate protection.

developmental biology

Mutations In Disordered Regions Cause Disease By Creating Endocytosis Motifs

Mutations in intrinsically disordered regions (IDRs) of proteins can cause a wide spectrum of diseases. Since IDRs lack a fixed three-dimensional structure, the mechanism by which such mutations cause disease is often unknown. Here, we employ a proteomic screen to investigate the impact of mutations in IDRs on protein-protein interactions. We find that mutations in disordered cytosolic regions of three transmembrane proteins (GLUT1, ITPR1 and CACNA1H) lead to an increased binding of clathrins. In all three cases, the mutation creates a dileucine motif known to mediate clathrin-dependent trafficking. Follow-up experiments on GLUT1 (SLC2A1), a glucose transporter involved in GLUT1 deficiency syndrome, revealed that the mutated protein mislocalizes to intracellular compartments. A systematic analysis of other known disease-causing variants revealed a significant and specific overrepresentation of gained dileucine motifs in cytosolic tails of transmembrane proteins. Dileucine motif gains thus appear to be a recurrent cause of disease.

biochemistry

HOT or not: Examining the basis of high-occupancy target regions

High-occupancy target (HOT) regions are the segments of the genome with unusually high number of transcription factor binding sites. These regions are observed in multiple species and thought to have biological importance due to high transcription factor occupancy. Furthermore, they coincide with house-keeping gene promoters and the associated genes are stably expressed across multiple cell types. Despite these features, HOT regions are solemnly defined using ChIP-seq experiments and shown to lack canonical motifs for transcription factors that are thought to be bound there. Although, ChIP-seq experiments are the golden standard for finding genome-wide binding sites of a protein, they are not noise free. Here, we show that HOT regions are likely to be ChIP-seq artifacts and they are similar to previously proposed \"hyper-ChIPable\" regions. Using ChIP-seq data sets for knocked-out transcription factors, we demonstrate presence of false positive signals on HOT regions. We observe sequence characteristics and genomic features that are discriminatory of HOT regions, such as GC/CpG-rich k-mers and enrichment of RNA-DNA hybrids (R-loops) and DNA tertiary structures (G-quadruplex DNA). The artificial ChIP-seq enrichment on HOT regions could be associated to these discriminatory features. Furthermore, we propose strategies to deal with such artifacts for the future ChIP-seq studies.

genomics