bioRxiv Science⌕ Search

Biology subjects

Cornette, J.

Publications and source records attributed to Cornette, J..

2 recordsLinked to original sources

Geographically Biased Composition of NetMHCpan Training Datasets and Evaluation of MHC-Peptide Binding Prediction Accuracy on Novel Alleles

Bias in neural network model training datasets has been observed to decrease prediction accuracy for groups underrepresented in training data. Thus, investigating the composition of training datasets used in machine learning models with health-care applications is vital to ensure equity. Two such machine learning models are NetMHCpan-4.1 and NetMHCIIpan-4.0, used to predict antigen binding scores to major histocompatibility complex class I and II molecules, respectively. As antigen presentation is a critical step in mounting the adaptive immune response, previous work has used these or similar predictions models in a broad array of applications, from explaining asymptomatic viral infection to cancer neoantigen prediction. However, these models have also been shown to be biased toward hydrophobic peptides, suggesting the network could also contain other sources of bias. Here, we report the composition of the networks training datasets are heavily biased toward European Caucasian individuals and against Asian and Pacific Islander individuals. We test the ability of NetMHCpan-4.1 and NetMHCpan-4.0 to distinguish true binders from randomly generated peptides on alleles not included in the training datasets. Unexpectedly, we fail to find evidence that the disparities in training data lead to a meaningful difference in prediction quality for alleles not present in the training data. We attempt to explain this result by mapping the HLA sequence space to determine the sequence diversity of the training dataset. Furthermore, we link the residues which have the greatest impact on NetMHCpan predictions to structural features for three alleles (HLA-A*34:01, HLA-C*04:03, HLA-DRB1*12:02).

immunology↗

Evasive spike variants elucidate the preservation of T cell immune response to the SARS-CoV-2 omicron variant

The Omicron variants boast the highest infectivity rates among all SARS-CoV-2 variants. Despite their lower disease severity, they can reinfect COVID-19 patients and infect vaccinated individuals as well. The high number of mutations in these variants render them resistant to antibodies that otherwise neutralize the spike protein of the original SARS-CoV-2 spike protein. Recent research has shown that despite its strong immune evasion, Omicron still induces strong T Cell responses similar to the original variant. This work investigates the molecular basis for this observation using the neural network tools NetMHCpan-4.1 and NetMHCiipan-4.0. The antigens presented through the MHC Class I and Class II pathways from all the notable SARS-CoV-2 variants were compared across numerous high frequency HLAs. All variants were observed to have equivalent T cell antigenicity. A novel positive control system was engineered in the form of spike variants that did evade T Cell responses, unlike Omicron. These evasive spike proteins were used to statistically confirm that the Omicron variants did not exhibit lower antigenicity in the MHC pathways. These results suggest that T Cell immunity mounts a strong defense against COVID-19 which is difficult for SARS-CoV-2 to overcome through mere evolution. Author summary

immunology↗