bioRxiv · 10.1101/2022.05.31.494108
CPPVec: an accurate coding potential predictor based on adistributed representation of protein sequence
Abstract
Long non-coding RNAs (lncRNAs) play a crucial role in numbers of biological processes and have received wide attention during the past years. Meanwhile, the rapid development of high-throughput transcriptome sequencing technologies (RNA-seq) lead to a large amount of RNA data, it is urgent to develop a fast and accurate coding potential predictor. Many computational methods have been proposed to alleviate this issue, they usually exploit information on open reading frame (ORF), k-mer, evolutionary signatures, or known protein databases. Despite the effectiveness, these methods still have much room to improve. Indeed, none of these methods exploit the context information of sequence, simple measures that are calculated with the continuous nucleotides are not enough to reflect global sequence order information. In view of this shortcoming, here, we present a novel alignment-free method, CPPVec, which exploits the global sequence order information of transcript for coding potential prediction for the first time, it can be easily implemented by distributed representation (e.g., doc2vec) of protein sequence translated from the longest ORF. Tests on human, mouse, zebrafish, fruit fly and Saccharomyces cerevisiae datasets demonstrate that CPPVec is an accurate coding potential predictor and significantly outperforms existing state-of-the-art methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wei, C., Ye, Z., Zhang, J., Li, A.. 2022-06-01. CPPVec: an accurate coding potential predictor based on adistributed representation of protein sequence. https://doi.org/10.1101/2022.05.31.494108
Cite the original work for its findings. Save a collection to share your selection of sources.