bioRxiv ScienceSearch

Biology subjects

Eyras, E.

Publications and source records attributed to Eyras, E..

3 recordsLinked to original sources

Combined analysis of genome sequencing and RNA-motifs reveals novel damaging non-coding mutations in human tumors

A major challenge in cancer research is to determine the biological and clinical significance of somatic mutations in non-coding regions. This has been studied in terms of recurrence, functional impact, and association to individual regulatory sites, but the combinatorial contribution of mutations to common RNA regulatory motifs has not been explored. We developed a new method, MIRA, to perform the first comprehensive study of significantly mutated regions (SMRs) affecting binding sites for RNA-binding proteins (RBPs) in cancer. Extracting signals related to RNA-related selection processes and using RNA sequencing data from the same samples we identified alterations in RNA expression and splicing linked to mutations on RBP binding sites. We found SRSF10 and MBNL1 motifs in introns, HNRPLL motifs at 5 UTRs, as well as 5 and 3 splice-site motifs, among others, with specific mutational patterns that disrupt the motif and impact RNA processing. MIRA facilitates the integrative analysis of multiple genome sites that operate collectively through common RBPs and can aid in the interpretation of non-coding variants in cancer. MIRA is available at https://github.com/comprna/mira.

cancer biology

SARNAclust: Semi-Automatic Detection Of RNA Protein Binding Motifs From Immunoprecipitation Data

RNA-protein binding is critical to gene regulation, controlling fundamental processes including splicing, translation, localization and stability, and aberrant RNA-protein interactions are known to play a role in a wide variety of diseases. However, molecular understanding of RNA-protein interactions remains limited, and in particular identification of the RNA motifs that bind proteins has long been a difficult problem. To address this challenge, we have developed a novel semi-automatic algorithm, SARNAclust, to computationally identify combined structure/sequence motifs from immunoprecipitation data. SARNAclust is, to our knowledge, the first unsupervised method that can identify RNA motifs at full structural resolution while also being able to simultaneously deconvolve multiple motifs. SARNAclust makes use of a graph kernel to evaluate similarity between sequence/structure objects, and provides the ability to isolate the impact of specific features through the bulge graph formalism. SARNAclust includes a key method for predicting RNA secondary structure at CLIP peaks, RNApeakFold, which we have verified to be effective on synthetic motif data. We applied SARNAclust to 30 ENCODE eCLIP datasets, identifying known motifs and novel predictions. Notably, we predicted a new motif for the protein ILF3 similar to that for the splicing factor hnRNPC, providing evidence for interaction between these two proteins. To validate our predictions, we performed a directed RNA bind-n-seq assay for two proteins: ILF3 and SLBP, in each case revealing the effectiveness of SARNAclust in predicting RNA sequence and structure elements important to protein binding. Availability: https://github.com/idotu/SARNAclust

bioinformatics

Fast and accurate differential splicing analysis across multiple conditions with replicates

Multiple approaches have been proposed to study differential splicing from RNA sequencing (RNA-seq) data1, including the analysis of transcript isoforms2,3, clusters of splice-junctions4,5, alternative splicing events6-8 and exonic regions9. However, many challenges remain unsolved, including the limitation in speed, the computing capacity and storage requirements, the constraints in the number of reads needed to achieve sufficient accuracy, and the lack of robust methods to account for variability between replicates and for analyses across multiple conditions. We present here a significant extension of SUPPA8 to enable streamlined analysis of differential splicing across multiple conditions, taking into account biological variability. We show that SUPPA differential splicing analysis achieves high accuracy using extensive experimental and simulated data compared to other methods; and shows higher accuracy at low sequencing depth, with short read lengths, and using replicas with unbalanced depth, which has important implications for the cost-effective use of RNA-seq data for splicing analysis. We also validate the analysis of multiple conditions with SUPPA by studying differential splicing during iPS-cell to neuron differentiation and during erythroblast differentiation, providing support for the applicability of SUPPA for the robust analysis of differential splicing beyond binary comparisons.

bioinformatics