bioRxiv Science⌕ Search

Biology subjects

Maquedano, M.

Publications and source records attributed to Maquedano, M..

3 recordsLinked to original sources

DEVELOPMENT OF A CONSENSUS MOLECULAR CLASSIFIER FOR PANCREATIC DUCTAL ADENOCARCINOMA

Pancreatic ductal adenocarcinoma (PDAC) presents a significant challenge, with a five-year survival rate of approximately 10%. Tumor heterogeneity contributes to the limited effectiveness of treatments. Several tumor and stroma molecular classifiers have attempted to clarify this heterogeneity with moderate agreement. Recognizing the complexity introduced by this extensive array of taxonomies, this study aims to develop a consensus molecular classifier by including both tumor and stroma features. We integrated gene expression data through Virtual Microdissection and classified the training samples to apply Machine Learning algorithms for each previous classifier. The consensus classifier was then derived using a Markov Clustering Algorithm, and its association with overall survival was assessed. The results indicated that Elastic-Net emerged as the superior model. We identified two classes for tumor components (Consensus Classical and Consensus Non-classical) and stroma components (Consensus Normal-Immune and Consensus Activated-ECM). The consensus Random Forest achieved a balanced accuracy of 96.33% and 98.92%, respectively. While not consistent across retrospective series, the algorithm (PDAConsensus) independently predicted overall survival. We developed a robust consensus classifier for PDAC that integrates tumor and stroma features and made it accessible through the R package PDACMOC (PDACMolecularOmniClassifier, https://github.com/pavillos/PDACMOC) and a Shiny app (https://pdacmoc.cnio.es/).

bioinformatics↗

More than 2,500 coding genes in the human reference gene set still have unsettled status

In 2018 we analysed the three main repositories for the human proteome, Ensembl/GENCODE, RefSeq and UniProtKB. They disagreed on the coding status of one of every eight annotated coding genes. The analysis inspired bilateral collaborations between annotation groups. Here we have repeated our analysis with updated versions of the three reference coding gene sets. Superficially, little appears to have changed. Although there are slightly fewer genes predicted as coding overall, the three groups still disagree on the status of 2,606 annotated genes. However, a comparison without read-through genes and immunoglobulin fragments shows that the three reference sets have merged or reclassified more than 700 genes since the last analysis and that just 0.6% of Ensembl/GENCODE coding genes are not also annotated by the other two reference sets. We used eight features indicative of non-coding genes to examine the 21,873 coding genes annotated across the three reference sets. We found that more than 2,000 had one or more potential non-coding features. While some of these genes will be protein coding, we believe that most are likely to be non-coding genes or pseudogenes. Our results suggest that annotators still vastly overestimate the number of true coding genes.

genomics↗

A deep audit of the PeptideAtlas database uncovers evidence for unannotated coding genes and aberrant translation

The human genome has been the subject of intense scrutiny by experimental and manual curation projects for more than two decades. Novel coding genes have been proposed from large-scale RNASeq, ribosome profiling and proteomics experiments. Here we carry out an in-depth analysis of an entire proteomics database. We analysed the proteins, peptides and spectra housed in the human build of the PeptideAtlas proteomics database to identify coding regions that are not yet annotated in the GENCODE reference gene set. We find support for hundreds of missing alternative protein isoforms and unannotated upstream translations, and evidence of cross-contamination from other species. There was reliable peptide evidence for 34 novel unannotated open reading frames (ORFs) in PeptideAtlas. We find that almost half belong to coding genes that are missing from GENCODE and other reference sets. Most of the remaining ORFs were not conserved beyond human, however, and their peptide confirmation was restricted to cancer cell lines. We show that this is strong evidence for aberrant translation, raising important questions about the extent of aberrant translation and how these ORFs should be annotated in reference genomes.

genomics↗