bioRxiv · 10.1101/671263
Recycling RNA-Seq Data to Identify Candidate Orphan Genes for experimental analysis
Abstract
The "dark transcriptome" can be considered the multitude of sequences that are transcribed but not annotated as genes. We evaluated expression of 6,692 annotated genes and 29,354 unannotated ORFs in the Saccharomyces cerevisiae genome across diverse environmental, genetic and developmental conditions (3,457 RNA-Seq samples). Over 48% of the transcribed ORFs have translation evidence. Phylostratigraphic analysis infers most of these transcribed ORFs would encode species-specific proteins ("orphan-ORFs"); hundreds have mean expression comparable to annotated genes. These data reveal unannotated ORFs most likely to be protein-coding genes. We partitioned a co-expression matrix by Markov Chain Clustering; the resultant clusters contain 2,468 orphan-ORFs. We provide the aggregated RNA-Seq yeast data with extensive metadata as a project in MetaOmGraph, a tool designed for interactive analysis and visualization. This approach enables reuse of public RNA-Seq data for exploratory discovery, providing a rich context for experimentalists to make novel, experimentally-testable hypotheses about candidate genes.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Li, J., Arendsee, Z., Singh, U., Wurtele, E. S.. 2019-06-21. Recycling RNA-Seq Data to Identify Candidate Orphan Genes for experimental analysis. https://doi.org/10.1101/671263
Cite the original work for its findings. Save a collection to share your selection of sources.