bioRxiv Science⌕ Search

Biology subjects

Kapil, S.

Publications and source records attributed to Kapil, S..

2 recordsLinked to original sources

Protein Language Model for Prediction of Subcellular Localization of Protein Sequences from Gram-negative bacteria (ProtLM.SCL)

The prediction of bacterial protein Sub-Cellular Localization (SCL) is critical for antigen identification and reverse vaccinology, especially when determining protein localization in the lab is time consuming, expensive and not possible for all species. While PSORTb is one of the most widely used tool for predicting SCL, it has several limitations, including the tendency to label a large number of proteins as Unknown. To address these shortcomings, we present a protein language model capable of predicting the subcellular localization of a given protein (ProtLM.SCL) from gram-negative bacteria. By performing 10-fold cross validation on the PSORTb public data set, we demonstrate that ProtLM.SCL is more accurate and precise than PSORTb. When compared to empirically validated published data, our models also outperformed PSORTb, particularly when categorizing difficult occurrences.

bioinformatics↗

T-cell receptor specific protein language model for prediction and interpretation of epitope binding (ProtLM.TCR)

The cellular adaptive immune response relies on epitope recognition by T-cell receptors (TCRs). We used a language model for TCRs (ProtLM.TCR) to predict TCR-epitope binding. This model was pre-trained on a large set of TCR sequences (~62.106) before being fine-tuned to predict TCR-epitope bindings across multiple human leukocyte antigen (HLA) of class-I types. We then tested ProtLM.TCR on a balanced set of binders and non-binders for each epitope, avoiding model shortcuts like HLA categories. We compared pan-HLA versus HLA-specific models, and our results show that while computational prediction of novel TCR-epitope binding probability is feasible, more epitopes and diverse training datasets are required to achieve a better generalized performances in de novo epitope binding prediction tasks. We also show that ProtLM.TCR embeddings outperform BLOSUM scores and hand-crafted embeddings. Finally, we have used the LIME framework to examine the interpretability of these predictions.

bioinformatics↗