bioRxiv Science⌕ Search

Biology subjects

Pesquita, C.

Publications and source records attributed to Pesquita, C..

2 recordsLinked to original sources

A comprehensive library of canonical and non-canonical MHC class I antigens for cancer vaccine development.

A longstanding disconnect between the growing number of MHC Class I immunopeptidomic studies and genomic medicine hinders cancer vaccine design. We develop COD-dipp to genomically map the full spectrum of detected canonical and non-canonical (non-exonic) MHC Class I antigens from 26 cancer studies. We demonstrate that patient mutations in regions overlapping physically identified antigens better predict immunotherapy response when compared to neoantigen predictions. We suggest a vaccine design approach using 140,966 highly immune-visible regions of the genome annotated by their expression and haplotype frequency in the human population. These regions tend to be highly conserved, mutated in cancer and harbor 7.8 times more immunogenicity. Intersecting pan-cancer mutations with these immune surveilled regions revealed a potential to create off-the-shelf multi-epitope vaccines against public neoantigens. Here we release COD-dipp, a cancer vaccine toolkit as a web-application (https://www.proteogenomics.ca/COD-dipp) and open-source high-throughput resource.

bioinformatics↗

Supervised biomedical semantic similarity

BackgroundSemantic similarity between concepts in knowledge graphs is essential for several bioinformatics applications, including the prediction of protein-protein interactions and the discovery of associations between diseases and genes. Although knowledge graphs describe entities in terms of several perspectives (or semantic aspects), state-of-the-art semantic similarity measures are general-purpose. This can represent a challenge since different use cases for the application of semantic similarity may need different similarity perspectives and ultimately depend on expert knowledge for manual fine-tuning. ResultsWe present a new approach that uses supervised machine learning to tailor aspect-oriented semantic similarity measures to fit a particular view on biological similarity or relatedness. We implement and evaluate it using different combinations of representative semantic similarity measures and machine learning methods with four biological similarity views: protein-protein interaction, protein function similarity, protein sequence similarity and phenotype-based gene similarity. ConclusionsThe results demonstrate that our approach outperforms non-supervised methods, producing semantic similarity models that fit different biological perspectives significantly better than the commonly used manual combinations of semantic aspects.

bioinformatics↗