Estimating Attentional Set-Shifting Dynamics in Varying Contextual Bandits
In this paper, we aim at estimating, on a trial-by-trial basis, the underlying decision-making process of an animal in a complex and changing environment. We propose a method for identifying the set of stochastic policies employed by the agent and estimating the transition dynamics between policies based on its behavior in a multidimensional discrimination task for measuring the properties of attentional set-shifting of the subject (both intra- and extra-dimensional). We propose using the Non-Homogeneous Hidden Markov Models (NHMMs) framework, to consider environmental state and rewards for modeling decision-making processes in a varying version of \"Contextual Bandits\". We employ the Expectation-Maximization (EM) procedure for estimating the models parameters similar to the Baum-Welch algorithm used to train standard HMMs. To measure the model capacity to estimate underlying dynamics, Monte Carlo analysis is employed on synthetically generated data and compared to the performance of classical HMM.