Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines
This paper develops optimization and Machine Learning (ML) algorithms to analyze gene expression datasets from the lungs and spleen of mice, infected intranasally, with two bacterial strains, Francisella tularensis - Schu4 and Live Vaccine Strain (LVS). We propose and utilize Weighted[l] 1-norm Generalized Eigenvalue-type Problems ([l]1-WGEPs) to determine a small set of host biomarkers that report Schu4 and LVS infection of the lungs and dissemination to the spleen. The optimal solutions of[l] 1-WGEPs determine the direction onto which the datasets are projected for dimensionality reduction, with the projection scores computed and ranked for gene selection. The top k-ranked projection scores correspond to the top k most informative biomarker features. The top k features selected from the lungs data are employed to train ML models, with uninfected controls and Schu4 or LVS samples as classes. The trained models are validated on the spleen data to incorporate transfer learning. Baseline ML algorithms such as ANN, XGBoost, AdaBoost, AdaGrad, KNN, SVM, Naive Bayes, Random Forest, Logistic Regression, and Decision Tree are compared with our Weighted[l] 1-norm Non-Parallel Proximal Support Vector Machine ([l]1-WNPSVM) that is based on two non-parallel separating hyperplanes. We report average balanced accuracy scores of the methods over multiple folds. Gene ontology is performed on the most significant genes in both tissues to reveal biomarkers of disease and examine for relevant metabolic pathways for host-directed therapeutics development and treatment performance. Author SummaryIntegrating genomic datasets from homogeneous or heterogeneous sources is an area that is currently underexplored. This work develops new methodologies to integrate transcriptomic datasets from the lungs and spleen tissues infected by Francisella tularensis -- Schu4 and Live Vaccine Strain (LVS). Our objective is to identify biologically relevant gene features indicative of respiratory infection, disease severity, and bacterial dissemination to the spleen, then utilize the selected features to predict disease status using our Weighted[l] 1-norm Non-Parallel Support Vector Machines ([l]1-WNPSVM), which is trained on the lungs data and validated on the spleen data, introducing a form of transfer learning. The[l] 1-WNPSVM outperforms traditional ML techniques, achieving a 97% balanced accuracy. It also generalizes to models of similar formulations, incorporating dimensionality reduction and gene selection into the NPSVM-type framework. Currently, a direct application of existing NPSVM-type methods to analyze gene expression datasets, where the number of genes significantly exceeds the number of samples, is computationally impractical due to their large memory requirements. This work addresses this challenge. We discovered sets of 253 genes exclusively expressed in the lungs and spleen tissues. Gene ontology is performed to reveal underlying metabolic pathways. Our analysis shows that the immune system pathway is activated in both lungs and spleen.