bioRxiv Science⌕ Search

Biology subjects

Anaya, J.

Publications and source records attributed to Anaya, J..

4 recordsLinked to original sources

DeepTMB: An uncertainty-aware deep calibration of tumor mutational burden with a synthetic tumor-only dataset

BackgroundTumor mutational burden (TMB) has been investigated as a biomarker for immune checkpoint blockade (ICB) therapy. Increasingly, TMB is being estimated with gene panel-based assays (as opposed to full exome sequencing) and different gene panels cover overlapping but distinct genomic coordinates, making comparisons across panels difficult. Previous studies have suggested that standardization and calibration to exome-derived TMB be done for each panel to ensure comparability. With TMB cutoffs being developed from panel-based assays, there is a need to understand how to properly estimate exomic TMB values from different panel-based assays. Design: Our approach to calibration of panel-derived TMB to exomic TMB proposes the use of probabilistic mixture models that allow for nonlinear relationships along with heteroscedastic error. We examined various inputs including nonsynonymous, synonymous, and hotspot counts along with genetic ancestry. Using the TCGA cohort we generated a tumor-only version of the panel-restricted data by reintroducing private germline variants. Results: We were able to model more accurately the distribution of both tumor-normal and tumor-only data using the proposed probabilistic mixture models as compared to linear regression. Applying a model trained on tumor-normal data to tumor-only input results in biased TMB predictions. Including synonymous mutations resulted in better regression metrics across both data types, but ultimately a model able to dynamically weight the various input mutation types exhibited optimal performance. Including genetic ancestry improved model performance only in the context of tumor-only data, wherein private germline variants are observed. SignificanceA probabilistic mixture model better models the nonlinearity and heteroscedasticity of the data as compared to linear regression. Tumor-only panel data is needed to properly calibrate tumor-only panels to exomic TMB. Leveraging the uncertainty of point estimates from these models better informs cohort stratification in terms of TMB.

bioinformatics↗

Read depth correction for somatic mutations

The ability to accurately detect mutations is a function of read depth and variant allele frequency (VAF). While the read depth distribution of a sample is observable, the true VAF distribution of all mutations in a sample is uncertain when there is low coverage depth. We propose to estimate the VAF distributions that would be observed with high-depth sequencing for samples with low sequencing depth by grouping samples with similar clonality and purity and using the VAF distributions observed with the high-depth mutations that are available. With these estimated high-depth VAF distributions we then calculate what the expected VAF distributions would be at a given depth and compare against the observed VAF distributions at that depth. Using this procedure we estimate that The Cancer Genome Atlas (TCGA) MC3 dataset only reports on average 83% of the mutations in a sample which would have been detected with high-depth sequencing. These results have important implications for comparing tumor mutational burden (TMB) estimates when samples are sequenced at different depths and for modeling high-depth, gene panel-based sequencing from the TCGA MC3 dataset.

bioinformatics↗

The proposed promiscuity value of an HLA can vary significantly depending on the source data used

Immune checkpoint blockade, a form of immunotherapy, mobilizes a patients own immune system against cancer cells by releasing some of the natural brakes on T cells. Although our understanding of this process is evolving, it is thought that a patient response to immunotherapy requires tumor presentation of neoantigens to T cells and patients whose tumors present a wider array of neoantigens are more likely to derive benefit from immune checkpoint blockade1-4. Manczinger et al.5 recently reported findings that would appear contrarian to this notion in that they suggested patients with HLA alleles which bind more diverse peptides (higher promiscuity) are less likely to respond to immunotherapy. To estimate HLA promiscuity they looked at the HLA-peptide binding repertoires for class I alleles contained in the IEDB6, and obtained consistent results when performing robustness checks and subsequent analyses. Here we show that the proposed HLA promiscuity values can vary significantly across source data types and individual experiments.

cancer biology↗

Aggregation Tool for Genomic Concepts (ATGC): A deep learning framework for sparse genomic measures and its application to tumor mutational burden

Deep learning can extract meaningful features from data given enough training examples. Large-scale genomic data are well suited for this class of machine learning algorithms; however, for many of these data the labels are at the level of the sample instead of at the level of the individual genomic measures. Conventional approaches to this data statically featurise and aggregate the measures separately from prediction. We propose to featurise, aggregate, and predict with a single trainable end-to-end model by turning to attention-based multiple instance learning. This allows for direct modelling of instance importance to sample-level classification in addition to trainable encoding strategies of genomic descriptions, such as mutations. We first demonstrate this approach by successfully solving synthetic tasks conventional approaches fail. Subsequently we applied the approach to somatic variants and achieved best-in-class performance when classifying tumour type or microsatellite status, while simultaneously providing an improved level of model explainability. Our results suggest that this framework could lead to biological insights and improve performance on tasks that aggregate information from sets of genomic data.

bioinformatics↗