Prediction and evaluation of Split-ORFs using Ribo-seq data
Split Open Reading frames (Split-ORFs) occur in transcripts containing at least two open reading frames, each encoding a part of the same full-length protein. These multiple open reading frames arise from alternatively spliced transcript isoforms. Understanding which genes make Split-ORFs, and in which cell types and under which conditions, would generate new insights into gene regulation. We previously published the Split-ORF pipeline, a computational tool that predicts candidate Split-ORFs from transcript sequences. Here, we present a new and improved version of the Split-ORF pipeline adding modules to analyze Ribo-seq data, calculate regions unique to the Split-ORF candidates, quantitatively assess Ribo-seq coverage in these regions, and perform candidate prioritization. Using this pipeline, we predicted more than 14,000 candidate Split-ORF transcripts from alternatively spliced human transcripts containing premature termination codons or retained introns. Hundreds of candidate Split-ORFs show significant Ribo-seq coverage across diverse cell types and diseases in at least one of the Split-ORFs, and 120 transcripts in both Split-ORFs. The candidate Split-ORF genes with significant Ribo-seq coverage are enriched for RNA-binding and RNA-processing functions and the majority of them encode RNA-binding proteins.