Potentials of Machine Learning in Predicting Key Features of Synthetic Antimicrobial Polymers (SAMPs)
The widespread antimicrobial resistance urged the need for novel antimicrobial agents. Synthetic antimicrobial polymers (SAMPs) were proposed as promising antibiotics to overcome the drawbacks of host-defence peptides. A machine learning forecast can be beneficial to evaluate the influence of SAMPs features on their potency and toxicity. In this study, we utilised a library of 20 polyacrylamides varied in: 1) type of amine side chain, 2) chain length, 3) cationic amine ratio, and 4) polymer architecture. Their structure-activity relationship was evaluated by comparing experimental observations and machine learning models. While classification models showed good fit in the training set, regression models demonstrated better fit in the testing set. Regression random forest and gradient boosting methods demonstrated reliable reproducibility of feature importance and maintained the tree structure throughout multiple runs. Beeswarm and waterfall plots provided an overview of the joint SHapley Additive exPlanations (SHAP) values of features on a specific data point. Based on validation tests and the consistency of feature importance and SHAP values, we conclude that boosting-ensemble methods can be utilized in forecasting future SAMPs.