bioRxiv ScienceSearch

Biology subjects

Nellore, A.

Publications and source records attributed to Nellore, A..

4 recordsLinked to original sources

neoepiscope Improves Neoepitope Prediction with Multi-variant Phasing

The vast majority of tools for neoepitope prediction from DNA sequencing of complementary tumor and normal patient samples do not consider germline context or the potential for co-occurrence of two or more somatic variants on the same mRNA transcript. Without consideration of these phenomena, existing approaches are likely to produce both false positive and false negative results, resulting in an inaccurate and incomplete picture of the cancer neoepitope landscape. We developed neoepiscope chiefly to address this issue for single nucleotide variants (SNVs) and insertions/deletions (indels), and herein illustrate how germline and somatic variant phasing affects neoepitope prediction across multiple datasets. We estimate that up to [~]5% of neoepitopes arising from SNVs and indels may require variant phasing for their accurate assessment. neoepiscope is performant, flexible, and supports several major histocompatibility complex binding affinity prediction tools. We have released neoepiscope as open-source software (MIT license, https://github.com/pdxgx/neoepiscope) for broad use.\n\nKEY POINTSO_LIGermline context and somatic variant phasing are important for neoepitope prediction\nC_LIO_LIMany popular neoepitope prediction tools have issues of performance and reproducibility\nC_LIO_LIWe describe and provide performant software for accurate neoepitope prediction from DNA-seq data\nC_LI

bioinformatics

RNA-seq transcript quantification from reduced-representation data in recount2

More than 70,000 short-read RNA-sequencing samples are publicly available through the recount2 project, a curated database of summary coverage data. However, no current methods can be directly applied to the reduced-representation information stored in this database to estimate transcript-level abundances. Here we present a linear model taking as input summary coverage of junctions and subdivided exons to output estimated abundances and associated uncertainty. We evaluate the performance of our model on simulated and real data, and provide a procedure to construct confidence intervals for estimates.

genomics

Population-level distribution and putative immunogenicity of cancer neoepitopes

BackgroundTumor neoantigens are a driver of cancer immunotherapy response; however, current neoantigen prediction tools produce many candidates that require further prioritization for research/clinical applications. Additional filtration criteria and population-level understanding may help to produce refined lists of putative neoantigens. Herein, we show neoepitope immunogenicity is likely related to measures of peptide novelty and report population-level behavior of these and other metrics.\n\nMethodsWe propose four peptide novelty metrics to refine predicted neoantigenicity: tumor vs. paired normal peptide binding affinity difference, tumor vs. paired normal peptide sequence similarity, tumor vs. closest human peptide sequence similarity, and tumor vs. closest microbial peptide sequence similarity. We apply these metrics to tumor neoepitopes predicted from somatic missense mutations in The Cancer Genome Atlas (TCGA) and a cohort of melanoma patients, as well as to a group of peptides with neoepitope-specific immune response data using an extension of pVAC-Seq [1].\n\nResultsWe show neoepitope burden varies across TCGA disease sites and HLA alleles, with surprisingly low repetition of neoepitope sequences across patients or neoepitope preferences among sets of HLA alleles. Only 20.3% of predicted neoepitopes across TCGA patients displayed novel binding change based on our binding affinity difference criteria. Similarity of amino acid sequence was typically high between paired tumor-normal epitopes, but in 24.6% of cases, neoepitopes were more similar to other human peptides, or even to bacterial (56.8% of cases) or viral peptides (15.5% of cases), than their paired normal counterparts. Applied to peptides with neoepitope-specific immune response, a linear model incorporating neoepitope binding affinity, protein sequence similarity between neoepitopes and their closest viral peptides, and paired binding affinity difference was able to predict immunogenicity with an AUROC of 0.66.\n\nConclusionsOur proposed neoepitope prioritization criteria emphasize neoepitope novelty and refine patient neoepitope predictions for focus on biologically meaningful candidate neoantigens. We have demonstrated that neoepitopes should be considered not only with respect to their paired normal epitope, but with respect to the entire human proteome, as well as bacterial and viral peptides, with potential implications for neoepitope immunogenicity and personalized vaccines for cancer treatment. We conclude that putative neoantigens are highly variable across individuals as a function of both cancer genetics and personalized HLA repertoire, while the overall behavior of filtration criteria reflects predictable patterns.

cancer biology

Snaptron: querying and visualizing splicing across tens of thousands of RNA-seq samples

As more and larger genomics studies appear, there is a growing need for comprehensive and queryable cross-study summaries. Snaptron is a search engine for summarized RNA sequencing data with a query planner that leverages R-tree, B-tree and inverted indexing strategies to rapidly execute queries over 146 million exon-exon splice junctions from over 70,000 human RNA-seq samples. Queries can be tailored by constraining which junctions and samples to consider. Snaptron can also rank and score junctions according to tissue specificity or other criteria. Further, Snaptron can rank and score samples according to the relative frequency of different splicing patterns. We outline biological questions that can be explored with Snaptron queries, including a study of novel exons in annotated genes, of exonization of repetitive element loci, and of a recently discovered alternative transcription start site for the ALK gene. Web app and documentation are at http://snaptron.cs.jhu.edu. Source code is at https://github.com/ChristopherWilks/snaptron under the MIT license.

bioinformatics