bioRxiv Science⌕ Search

Biology subjects

Broadwell, Q.

Publications and source records attributed to Broadwell, Q..

2 recordsLinked to original sources

Pep2Vec: An Interpretable Model for Peptide-MHC Presentation Prediction and Contaminant Identification in Ligandome Datasets

As personalized cancer vaccines advance, precise modeling of antigen presentation by MHC class I and II is crucial. High-quality training data is essential for clinical models. Existing deep learning models focus on prediction performance but lack interpretability. We introduce Pep2Vec, a modular, transformer-based model trained on MHC I and II ligandome data, transforming input sequences into interpretable vectors. This approach integrates source protein features and elucidates the source of its performance gains, revealing regions that correlate with gene expression and protein-protein interactions. Pep2Vecs peptide latent space shows relationships between peptides of varying MHC class, allotype, lengths, and submotifs. This enables identifying four major contaminant types, constituting 5.0% of our data. Pep2Vec enhances MHC presentation prediction, achieving higher average precision on our presentation test set and immunogenicity datasets than existing models, and reducing contaminant-like peptide recommendations. Pep2Vec addresses a critical need for the development of more precise and effective applications of peptide MHC models, such as for cancer vaccines and antibody deimmunization.

immunology↗

HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features.

Antigen presentation on MHC class I (MHC-I) is key to the adaptive immune response to cancerous cells. Computational prediction of peptide presentation by MHC-I has enabled individualized cancer immunotherapies. Here, we introduce HLApollo, a transformer-based approach with end-to-end modeling of MHC-I sequence, deconvolution, and flanking sequences. To achieve this, we develop a novel training strategy, negative set switching, which greatly reduces overfitting to falsely presumed negatives that are necessarily found in presentation datasets. HLApollo shows a meaningful improvement compared to recent MHC-I models on peptide presentation (20.19% average precision (AP)) and immunogenicity (4.1% AP). As expected, adding gene expression boosts the performance of HLApollo. More interestingly, we show that introduction of features from a protein language model, ESM 1b, remarkably recoups much of the benefits of gene expression in absence of true expression measurements. Finally, we demonstrate excellent pan-allelic generalization, and introduce a framework for estimating the expected accuracy of HLApollo for untrained alleles. This guides the use of HLApollo in a clinical setting, where rare alleles may be observed in some subjects, particularly for underrepresented minorities.

immunology↗