bioRxiv ScienceSearch

Biology subjects

Uwe Ohler

Publications and source records attributed to Uwe Ohler.

3 recordsLinked to original sources

ssHMM: Extracting intuitive sequence-structure motifs from high-throughput RNA-binding protein data

RNA-binding proteins (RBPs) play important roles in RNA post-transcriptional regulation and recognize target RNAs via sequence-structure motifs. To which extent RNA structure influences protein binding in the presence or absence of a sequence motif is still poorly understood. Existing RNA motif finders which produce informative motifs and simultaneously capture the relationship between primary sequence and different RNA secondary structures are missing. We developed ssHMM, an RNA motif finder that combines a hidden Markov model (HMM) with Gibbs sampling to learn the joint sequence and structure binding preferences of RBPs from high-throughput data, such as CLIP-Seq sequences, and visualizes them as a graph. Evaluations on synthetic data showed that ssHMM reliably recovers fuzzy sequence motifs in 80 to 100% of the cases. It produces motifs with higher information content than existing tools and is faster than other methods on large datasets. Examples of new sequence-structure motifs identified by ssHMM for uncharacterized RBPs are also discussed. ssHMM is freely available on Github at https://github.molgen.mpg.de/heller/ssHMM.

Bioinformatics

Integrative classification of human coding and non-coding genes based on RNA metabolism profiles

The pervasive transcription of the human genome results in a heterogeneous mix of coding and long non-coding RNAs (lncRNAs). Only a small fraction of lncRNAs possess demonstrated regulatory functions, making it difficult to distinguish functional lncRNAs from non-functional transcriptional byproducts. This has resulted in numerous competing classifications of human lncRNA that are complicated by a steady increase in the number of annotated lncRNAs.\n\nTo address these challenges, we quantitatively examined transcription, splicing, degradation, localization and translation for coding and non-coding human genes. Annotated lncRNAs had lower synthesis and higher degradation rates than mRNAs, and we discovered mechanistic differences explaining the slower splicing of lncRNAs. We grouped genes into classes with similar RNA metabolism profiles. These classes contained both mRNAs and lncRNAs to varying degrees; they exhibited distinct relationships between steps of RNA metabolism, evolutionary patterns, and sensitivity to cellular RNA regulatory pathways. Our classification provides a behaviorally-coherent alternative to genomic context-driven annotations of lncRNAs.\n\nHighlightsO_LIHigh-resolution 4SU pulse labeling of RNA allows for quantifying synthesis, processing and decay rates across thousands of coding and non-coding transcripts.\nC_LIO_LISynthesis and processing rates of lncRNAs are lower than mRNAs, while degradation rates were substantially higher\nC_LIO_LIDifferences in the splicing efficiency between slow/lncRNA and fast/mRNA introns are explained by GC-content, splicing regulatory elements and unphosphorylated RNA poll II.\nC_LIO_LIA new annotation-agnostic classification of RNAs reveals seven clusters of lncRNAs and mRNAs with unique metabolism patterns that provides behaviorally coherent subsets of lncRNAs.\nC_LIO_LIClasses are distinguished by evolutionary patterns and sensitivity to cellular RNA regulatory pathways.\nC_LI

Genomics

A spectral analysis approach to detect actively translated open reading frames in high-resolution ribosome profiling data

RNA sequencing protocols allow for quantifying gene expression regulation at each individual step, from transcription to protein synthesis. Ribosome Profiling (Ribo-seq) maps the positions of translating ribosomes over the entire transcriptome. Despite its great potential, a rigorous statistical approach to identify translated regions by means of the characteristic three-nucleotide periodicity of Ribo-seq data is not yet available. To fill this gap, we developed RiboTaper, which quantifies the significance of periodic Ribo-seq reads via spectral analysis methods.\n\nWe applied RiboTaper on newly generated, deep Ribo-seq data in HEK293 cells, to derive an extensive map of translation that covers Open Reading Frame (ORF) annotations for more than 11,000 protein-coding genes. We also find distinct ribosomal signatures for several hundred detected upstream ORFs and ORFs in annotated non-coding genes (ncORFs). Mass spectrometry data confirms that RiboTaper achieves excellent coverage of the cellular proteome and validates dozens of novel peptide products. Collectively, RiboTaper (available at https://ohlerlab.mdc-berlin.de/software/) is a powerful method for comprehensive de novo identification of actively used ORFs in the human genome.

Genomics