bioRxiv Science⌕ Search

Biology subjects

Cerdan-Velez, D.

Publications and source records attributed to Cerdan-Velez, D..

4 recordsLinked to original sources

More than 100 dual coding regions have evidence for selection constraints in both reading frames

Alternative splicing can generate multiple differently spliced transcripts from a single pre-mRNA. A striking number of genes have alternative splice events that ccan hange the downstream reading frame leading to exons that code from distinct reading frames. In fact, more than a third of the coding genes in the human gene set are annotated with dual coding exons derived from alternative splicing events. Here we analysed a set of 537 dual coding regions that have evidence to support their functional importance. These dual coding regions produce protein isoforms with completely different C-terminals and have reading frames that are supported by either peptide or conservation evidence. More than a quarter of the alternative reading frames are preserved across all mammals, and many can be traced back to the earliest jawed vertebrates. Most of these ancient dual coding regions appear to be under selective constraints. We find support for purifying selection on both frames in 105 pairs of transcripts and two genes, CCSER2 and SH2B1, have triple coding regions that are under clear selection pressure in all three frames. We found evidence to suggest that many ancient dual coding regions may have played important roles in the evolution of the vertebrate central nervous system. Most ancient dual coding regions with evidence for protein level tissue specificity were brain specific and we showed that genes with ancient dual coding regions are highly enriched in brain tissues. Most remarkably, we found that more than 80% of the genes with these ancient dual coding regions are implicated in neuron development, synapses and neural cell projections.

genomics↗

More than 2,500 coding genes in the human reference gene set still have unsettled status

In 2018 we analysed the three main repositories for the human proteome, Ensembl/GENCODE, RefSeq and UniProtKB. They disagreed on the coding status of one of every eight annotated coding genes. The analysis inspired bilateral collaborations between annotation groups. Here we have repeated our analysis with updated versions of the three reference coding gene sets. Superficially, little appears to have changed. Although there are slightly fewer genes predicted as coding overall, the three groups still disagree on the status of 2,606 annotated genes. However, a comparison without read-through genes and immunoglobulin fragments shows that the three reference sets have merged or reclassified more than 700 genes since the last analysis and that just 0.6% of Ensembl/GENCODE coding genes are not also annotated by the other two reference sets. We used eight features indicative of non-coding genes to examine the 21,873 coding genes annotated across the three reference sets. We found that more than 2,000 had one or more potential non-coding features. While some of these genes will be protein coding, we believe that most are likely to be non-coding genes or pseudogenes. Our results suggest that annotators still vastly overestimate the number of true coding genes.

genomics↗

A deep audit of the PeptideAtlas database uncovers evidence for unannotated coding genes and aberrant translation

The human genome has been the subject of intense scrutiny by experimental and manual curation projects for more than two decades. Novel coding genes have been proposed from large-scale RNASeq, ribosome profiling and proteomics experiments. Here we carry out an in-depth analysis of an entire proteomics database. We analysed the proteins, peptides and spectra housed in the human build of the PeptideAtlas proteomics database to identify coding regions that are not yet annotated in the GENCODE reference gene set. We find support for hundreds of missing alternative protein isoforms and unannotated upstream translations, and evidence of cross-contamination from other species. There was reliable peptide evidence for 34 novel unannotated open reading frames (ORFs) in PeptideAtlas. We find that almost half belong to coding genes that are missing from GENCODE and other reference sets. Most of the remaining ORFs were not conserved beyond human, however, and their peptide confirmation was restricted to cancer cell lines. We show that this is strong evidence for aberrant translation, raising important questions about the extent of aberrant translation and how these ORFs should be annotated in reference genomes.

genomics↗

GENCODE: massively expanding the lncRNA catalog through capture long-read RNA sequencing

Accurate and complete gene annotations are indispensable for understanding how genome sequences encode biological functions. For more than twenty years, the GENCODE consortium has developed reference annotations for the human and mouse genomes, becoming a foundation for biomedical and genomics communities worldwide. Nevertheless, collections of important yet poorly-understood gene classes like long non-coding RNAs (lncRNAs) remain incomplete and scattered across multiple, uncoordinated catalogs. To address this, GENCODE has undertaken the most comprehensive lncRNA annotation effort to date. This is founded on the manually supervised computational annotation of full-length targeted long-read sequencing, on matched embryonic and adult tissues, of orthologous regions in human and mouse. Altogether 17,931 human genes (140,268 transcripts) and 22,784 mouse genes (136,169 transcripts) have been added to the GENCODE catalog representing a 2-fold and 6-fold growth in transcripts, respectively - the greatest increase in the number of annotated human genes since the sequencing of the human genome. Our targeted design assigned human-mouse orthologs at a rate beyond previous studies, tripling the number of human disease-associated lncRNAs that have mouse orthologs. Novel lncRNA genes consistently exhibit biological signals of functionality, and they greatly enhance the functional interpretability of the human genome. While poorly expressed in bulk RNA-Seq samples, many of them are highly expressed in specific cell populations, maybe even contributing to cell-type determination. The expanded GENCODE lncRNA annotations mark a critical step toward deciphering the human and mouse genomes.

genomics↗