bioRxiv ScienceSearch

Biology subjects

Backofen, R.

Publications and source records attributed to Backofen, R..

8 recordsLinked to original sources

uORF-Tools - Workflow for the determination of translation-regulatory upstream open reading frames

Ribosome profiling (ribo-seq) provides a means to analyze active translation by determining ribosome occupancy in a transcriptome-wide manner. The vast majority of ribosome protected fragments resides within the protein-coding sequence of mRNAs. However, commonly reads are also found within the transcript leader sequence (TLS) (aka 5 untranslated region) preceding the main open reading frame (ORF), which indicates the translation of regulatory upstream ORFs (uORFs). Here, we present a workflow for the identification of functional uORFs, which contribute to the translational regulation of their associated main ORFs. The workflow is available as free and open Snakemake workflow. Furthermore, we provide a comprehensive human uORF annotation file, which can be used within the pipeline, thus reducing the runtime. (Availability: https://github.com/anibunny12/uORF-Tools)

bioinformatics

Structure probing data enhances RNA-RNA interaction prediction

SummaryExperimental structure probing data has been shown to improve thermodynamics-based RNA secondary structure prediction. To this end, chemical reactivity information (as provided e.g. by SHAPE) is incorporated, which encodes whether or not individual nucleotides are involved in intra-molecular structure. Since inter-molecular RNA-RNA interactions are often confined to unpaired RNA regions, SHAPE data is even more promising to improve interaction prediction. Here we show how such experimental data can be incorporated seamlessly into accessibility-based RNA-RNA interaction prediction approaches, as implemented in IntaRNA. This is possible via the computation and use of unpaired probabilities that incorporate the structure probing information. We show that experimental SHAPE data can significantly improve RNA-RNA interaction prediction. We evaluate our approach by investigating interactions of a spliceosomal U1 snRNA transcript with its target splice sites. When SHAPE data is incorporated, known target sites are predicted with increased precision and specificity.\n\nAvailabilityhttps://github.com/BackofenLab/IntaRNA

bioinformatics

MechRNA: prediction of lncRNA mechanisms from RNA-RNA and RNA-protein interactions

MotivationLong non-coding RNAs (lncRNAs) are defined as transcripts longer than 200 nucleotides that do not get translated into proteins. Often these transcripts are processed (spliced, capped, polyadenylated) and some are known to have important biological functions. However, most lncRNAs have unknown or poorly understood functions. Nevertheless, because of their potential role in cancer, lncRNAs are receiving a lot of attention, and the need for computational tools to predict their possible mechanisms of action is more than ever. Fundamentally, most of the known lncRNA mechanisms involve RNA-RNA and/or RNA-protein interactions. Through accurate predictions of each kind of interaction and integration of these predictions, it is possible to elucidate potential mechanisms for a given lncRNA.\n\nApproachHere we introduce MechRNA, a pipeline for corroborating RNA-RNA interaction prediction and protein binding prediction for identifying possible lncRNA mechanisms involving specific targets or on a transcriptome-wide scale. The first stage uses a version of IntaRNA2 with added functionality for efficient prediction of RNA-RNA interactions with very long input sequences, allowing for large-scale analysis of lncRNA interactions with little or no loss of optimality. The second stage integrates protein binding information pre-computed by GraphProt, for both the lncRNA and the target. The final stage involves inferring the most likely mechanism for each lncRNA/target pair. This is achieved by generating candidate mechanisms from the predicted interactions, the relative locations of these interactions and correlation data, followed by selection of the most likely mechanistic explanation using a combined p-value.\n\nResultsWe applied MechRNA on a number of recently identified cancer-related lncRNAs (PCAT1, PCAT29, ARLnc1) and also on two well-studied lncRNAs (PCA3 and 7SL). This led to the identification of hundreds of high confidence potential targets for each lncRNA and corresponding mechanisms. These predictions include the known competitive mechanism of 7SL with HuR for binding on the tumor suppressor TP53, as well as mechanisms expanding what is known about PCAT1 and ARLn1 and their targets BRCA2 and AR, respectively. For PCAT1-BRCA2, the mechanism involves competitive binding with HuR, which we confirmed using HuR immunoprecipitation assays.\n\nAvailabilityMechRNA is available for download at https://bitbucket.org/compbio/mechrna\n\nContactbackofen@informatik.uni-freiburg.de, cenksahi@indiana.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Conserved accessory proteins encoded with archaeal and bacterial Type III CRISPR-Cas gene cassettes that may specifically modulate, complement or extend interference activity

A study was undertaken to identify conserved proteins that are encoded either within, or directly adjacent to, cas gene cassettes of Type III CRISPR-Cas interference modules. These Type III modules are especially versatile functionally and have been shown to target and degrade dsDNA, ssDNA and ssRNA. In addition, the interference gene cassettes are frequently intertwined with other accessory genes, including genes encoding CARF domains, some of which are likely to be cofunctional. Using a comparative genomics approach, and defining a Type III association score accounting for coevolution and specificity of flanking genes, we identified and classified 39 new Type III associated gene families. Most archaeal and bacterial Type III modules were seen to be flanked by several accessory genes, around half of which did not encode CARF domains and remain of unknown function. Non-CARF accessory genes were found to be more diverse than their CARF counterparts, encoding nuclease, helicase, protease, ATPase, transporter and transmembrane domains and including a considerable fraction that encoded no known domains. The diversity of non-CARF Type III accessory genes found in this study suggests that additional families exist which remain undetected because of the limited number of annotated genomes currently available. The method employed is scalable for potential application on metagenomic data once automated pipelines for annotation of CRISPR-Cas systems have been developed. All accessory genes found in this study are presented online in a readily accessible and searchable format for researchers to audit their model organism of choice: http://accessory.crispr.dk.

bioinformatics

Community-driven data analysis training for biology

The primary problem with the explosion of biomedical datasets is not the data itself, not computational resources, and not the required storage space, but the general lack of trained and skilled researchers to manipulate and analyze these data. Eliminating this problem requires development of comprehensive educational resources. Here we present a community-driven framework that enables modern, interactive teaching of data analytics in life sciences and facilitates the development of training materials. The key feature of our system is that it is not a static but a continuously improved collection of tutorials. By coupling tutorials with a web-based analysis framework, biomedical researchers can learn by performing computation themselves through a web-browser without the need to install software or search for example datasets. Our ultimate goal is to expand the breadth of training materials to include fundamental statistical and data science topics and to precipitate a complete re-engineering of undergraduate and graduate curricula in life sciences.

bioinformatics

Distinct epigenetic programs regulate cardiac myocyte development and disease in the human heart in vivo

Epigenetic mechanisms and transcription factor networks essential for differentiation of cardiac myocytes have been uncovered. However, reshaping of the epigenome of these terminally differentiated cells during fetal development, postnatal maturation and in disease remains unknown. Thus, the aim of this study was to determine the dynamics of the cardiac myocyte epigenome during development and in chronic heart failure. Prenatal development and postnatal maturation are characterized by a cooperation of active CpG methylation and histone marks at cis-regulatory and genic regions to shape the cardiac myocyte transcriptome. In contrast, pathological gene expression in terminal heart failure is accompanied by changes in active histone marks without major alterations in CpG methylation and repressive chromatin marks. Notably, cis-regulatory regions in cardiac myocytes are significantly enriched for cardiovascular disease-associated variants. This study uncovers distinct layers of epigenetic regulation not only during prenatal development and postnatal maturation but also in diseased human cardiac myocytes.

genomics

Practical computational reproducibility in the life sciences

Many areas of research suffer from poor reproducibility. This problem is particularly acute in computationally intensive domains where results rely on a series of complex methodological decisions that are not well captured by traditional publication approaches. Various guidelines have emerged for achieving reproducibility, but practical implementation of these practices remains difficult. This is because reproducing published computational analyses requires installing many software tools plus associated libraries, connecting tools together into the complete pipeline, and specifying parameters. Here we present a suite of recently emerged technologies which make computational reproducibility not just possible, but, finally, practical in both time and effort. By combining a system for building highly portable packages of bioinformatics software, containerization and virtualization technologies for isolating reusable execution environments for these packages, and an integrated workflow system that automatically orchestrates the composition of these packages for entire pipelines, an unprecedented level of computational reproducibility can be achieved.

bioinformatics

uvCLAP: a fast, non-radioactive method to identify in vivo targets of RNA-binding proteins

RNA-binding proteins (RBPs) play important and essential roles in eukaryotic gene expression regulating splicing, localization, translation and stability of mRNAs. Understanding the exact contribution of RBPs to gene regulation is crucial as many RBPs are frequently mis-regulated in several neurological diseases and certain cancers. While recently developed techniques provide binding sites of RBPs, they are labor-intensive and generally rely on radioactive labeling of RNA. With more than 1,000 RBPs in a human cell, it is imperative to develop easy, robust, reproducible and high-throughput methods to determine in vivo targets of RBPs. To address these issues we developed uvCLAP (UV crosslinking and affinity purification) as a robust, reproducible method to measure RNA-protein interactions in vivo. To test its performance and applicability we investigated binding of 15 RBPs from fly, mouse and human cells. We show that uvCLAP generates reliable and comparable data to other methods. Unexpectedly, our results show that despite their different subcellular localizations, STAR proteins (KHDRBS1-3, QKI) bind to a similar RNA motif in vivo. Consistently a point mutation (KHDRBS1Y440F) or a natural splice isoform (QKI-6) that changes the respective RBP subcellular localization, dramatically alters target selection without changing the targeted RNA motif. Combined with the knowledge that RBPs can compete and cooperate for binding sites, our data shows that compartmentalization of RBPs can be used as an elegant means to generate RNA target specificity.

molecular biology