bioRxiv Science⌕ Search

Biology subjects

Atkins, T. K.

Publications and source records attributed to Atkins, T. K..

2 recordsLinked to original sources

Geographically Biased Composition of NetMHCpan Training Datasets and Evaluation of MHC-Peptide Binding Prediction Accuracy on Novel Alleles

Bias in neural network model training datasets has been observed to decrease prediction accuracy for groups underrepresented in training data. Thus, investigating the composition of training datasets used in machine learning models with health-care applications is vital to ensure equity. Two such machine learning models are NetMHCpan-4.1 and NetMHCIIpan-4.0, used to predict antigen binding scores to major histocompatibility complex class I and II molecules, respectively. As antigen presentation is a critical step in mounting the adaptive immune response, previous work has used these or similar predictions models in a broad array of applications, from explaining asymptomatic viral infection to cancer neoantigen prediction. However, these models have also been shown to be biased toward hydrophobic peptides, suggesting the network could also contain other sources of bias. Here, we report the composition of the networks training datasets are heavily biased toward European Caucasian individuals and against Asian and Pacific Islander individuals. We test the ability of NetMHCpan-4.1 and NetMHCpan-4.0 to distinguish true binders from randomly generated peptides on alleles not included in the training datasets. Unexpectedly, we fail to find evidence that the disparities in training data lead to a meaningful difference in prediction quality for alleles not present in the training data. We attempt to explain this result by mapping the HLA sequence space to determine the sequence diversity of the training dataset. Furthermore, we link the residues which have the greatest impact on NetMHCpan predictions to structural features for three alleles (HLA-A*34:01, HLA-C*04:03, HLA-DRB1*12:02).

immunology↗

FIST-nD: A tool for n-dimensional spatial transcriptomics data imputation via graph-regularized tensor completion

Functional interpretation of spatial transcriptomics data usually requires non-trivial pre-processing steps and other supporting data in the analysis due to the high sparsity and incompleteness of spatial RNA profiling, especially in 3D constructions. As a solution, we present a new software tool FIST-nD, Fast Imputation of Spatially-resolved transcriptomes by graph-regularized Tensor completion in n-Dimensions for imputing 3D as well as 2D spatial transcriptomics data. FIST-nD is implemented based on a novel graph-regularized tensor decomposition method, which imputes spatial gene expression data using 4-way high-order tensor structure and relations in spatial and gene functional graphs. The implementation, accelerated by GPU or multicore parallel computing, can efficiently impute high-resolution 3D spatial transcriptomics data within a few minutes. The experiments on three 3D Spatial Transcriptomics datasets and one 3D high-resolution Stereo-seq dataset confirm the high accuracy of the imputation by FIST-nD and demonstrate that the imputed spatial transcriptomes provide a more complete gene expression landscape for downstream analyses such as spatial gene expression clustering and visualizations.

bioinformatics↗