bioRxiv ScienceSearch

Biology subjects

Dayan, P.

Publications and source records attributed to Dayan, P..

9 recordsLinked to original sources

Forgetful inference in a sophisticated world model

Humans and other animals are able to discover underlying statistical structure in their environments and exploit it to achieve efficient and effective performance. However, such structure is often difficult to learn and use because it is obscure, involving long-range temporal dependencies. Here, we analysed behavioural data from an extended experiment with rats, showing that the subjects learned the underlying statistical structure, albeit suffering at times from immediate inferential imperfections as to their current state within it. We accounted for their behaviour using a Hidden Markov Model, in which recent observations are integrated with the recollections of an imperfect memory. We found that over the course of training, subjects came to track their progress through the task more accurately, a change that our model largely attributed to decreased forgetting. This learning to remember decreased reliance on recent observations, which may be misleading, in favour of a longer-term memory.\n\nAuthor summaryHumans and other animals possess the remarkable ability to find and exploit patterns and structures in their experience of a complex and varied world. However, such structures are often temporally extended and latent or hidden, being only partially correlated with immediate observations of the world. This makes it essential to integrate current and historical information, and creates a challenging statistical and computational problem.\n\nHere, we examine the behaviour of rats facing a version of this challenge posed by a brain-stimulation reward task. We find that subjects learned the general structure of the task, but struggled when immediate observations were misleading. We captured this behaviour with a model in which subjects integrated evidence from their observations together with a memory whose imperfections accounted for their errors.\n\nThe subjects performance improved markedly over successive sessions, allowing them to overcome misleading observations. According to the model, this arose from a process of learning to remember in which subjects became better at employing more reliable past observations to determine the hidden state of the world.

neuroscience

Neural signatures of detours, shortcuts and back-tracking during navigation

Central to the concept of the cognitive map is that it confers behavioural flexibility, allowing animals to take efficient detours, exploit shortcuts and realise the need to back-track rather than persevere on a poorly chosen route. The neural underpinnings of such naturalistic and flexible behaviour remain unclear. During fMRI we tested human subjects on their ability to navigate to a set of goal locations in a virtual desert island riven by lava, which occasionally shifted to block selected paths (necessitating detours) or receded to open new paths (affording shortcuts). We found that during self-initiated back-tracking, activity increased in frontal regions and the dorsal anterior cingulate cortex, while activity in regions associated with the core default-mode network was suppressed. Detours activated a network of frontal regions compared to shortcuts. Activity in right dorsolateral prefrontal cortex specifically increased when participants encountered new plausible shortcuts but which in fact added to the path (false shortcuts). These results help inform current models as to how the brain supports navigation and planning in dynamic environments.\n\nSignificance StatementAdaptation to change is important for survival. Although real-world spatial environments are prone to continual change, little is known about how the brain supports navigation in dynamic environments where flexible adjustments to route plans are needed. Here, we used fMRI to examine the brain activity elicited when humans took forced detours, identified shortcuts and spontaneously back-tracked along their recent path. Both externally and internally generated changes in the route activated the fronto-parietal attention network, whereas only internally generated changes generated increased activity in the dorsal anterior cingulate cortex with a concomitant disengagement in regions associated with the default-mode network. The results provide new insights into how the brain plans and re-plans in the face of a changing environment.

neuroscience

Integrated accounts of behavioral and neuroimaging data using flexible recurrent neural network models

Neuroscience studies of human decision-making abilities commonly involve sub-jects completing a decision-making task while BOLD signals are recorded using fMRI. Hypotheses are tested about which brain regions mediate the effect of past experience, such as rewards, on future actions. One standard approach to this is model-based fMRI data analysis, in which a model is fitted to the behavioral data, i.e., a subjects choices, and then the neural data are parsed to find brain regions whose BOLD signals are related to the models internal signals. However, the internal mechanics of such purely behavioral models are not constrained by the neural data, and therefore might miss or mischaracterize aspects of the brain. To address this limitation, we introduce a new method using recurrent neural network models that are flexible enough to be jointly fitted to the behavioral and neural data. We trained a model so that its internal states were suitably related to neural activity during the task, while at the same time its output predicted the next action a subject would execute. We then used the fitted model to create a novel visualization of the relationship between the activity in brain regions at different times following a reward and the choices the subject subsequently made. Finally, we validated our method using a previously published dataset. We found that the model was able to recover the underlying neural substrates that were discovered by explicit model engineering in the previous work, and also derived new results regarding the temporal pattern of brain activity.

neuroscience

Spotting the path that leads nowhere: Modulation of human theta and alpha oscillations induced by trajectory changes during navigation

The capacity to take efficient detours and exploit novel shortcuts during navigation is thought to be supported by a cognitive map of the environment. Despite advances in understanding the neural basis of the cognitive map, little is known about the neural dynamics associated with detours and shortcuts. Here, we recorded magnetoencephalography from humans as they navigated a virtual desert island riven by shifting lava flows. The task probed their ability to take efficient detours and shortcuts to remembered goals. We report modulation in event-related fields and theta power as participants identified real shortcuts and differentiated these from false shortcuts that led along suboptimal paths. Additionally, we found that a decrease in alpha power preceded back-tracking where participants spontaneously turned back along a previous path. These findings help advance our understanding of the fine-grained temporal dynamics of human brain activity during navigation and support the development of models of brain networks that support navigation.

neuroscience

Models that learn how humans learn: the case of depression and bipolar disorders

Computational models of learning and decision-making processes in the brain play an important role in many domains. Such models typically have a constrained structure and make specific assumptions about the underlying human learning processes; these may make them underfit observed behaviours. Here we suggest an alternative method based on learning-to-learn approaches, using recurrent neural networks (RNNs) as a flexible family of models that have sufficient capacity to represent the complex learning and decision-making strategies used by humans. In this approach, an RNN is trained to predict the next action that a subject will take in a decision-making task, and in this way, learns to imitate the processes underlying subjects choices and their learning abilities. We demonstrate the benefits of this approach with a new dataset containing behaviour of uni-polar depression (n=34), bipolar (n=33) and control (n=34) participants in a two-armed bandit task. The results indicate that the new approach is better than baseline reinforcement-learning methods in terms of overall performance and its capacity to predict subjects choices. We show that the model can be interpreted using off-policy simulations, and thereby provide a novel clustering of subjects learning processes - something that often eludes traditional approaches to modelling and behavioural analysis.

bioinformatics

Noradrenaline modulates decision urgency during sequential information gathering

Arbitrating between timely choice and extended information gathering is critical in effective decision making. Aberrant information gathering behaviour is said to be a feature of psychiatric disorders such as schizophrenia and obsessive-compulsive disorder. We know little about the neurocognitive control mechanisms that drive such information gathering. In a double-blind placebo-controlled drug study with 60 healthy humans (30 female), we examined the effects of noradrenaline and dopamine antagonism on information gathering. We show that modulating noradrenaline function with propranolol leads to decreased information gathering behaviour and this contrasts with no effect following a modulation of dopamine function. Using a Bayesian computational model, we show sampling behaviour is best explained when including an urgency signal that promotes commitment to an early decision. We demonstrate that noradrenaline blockade promotes the expression of this decision-related urgency signal during information gathering. We discuss the findings with respect to psychopathological conditions that are linked to aberrant information gathering.\n\nSignificance StatementKnowing when to stop gathering information and commit to an option is non-trivial. This is an important element in arbitrating between information gain and energy conservation. In this double-blind, placebo-controlled drug study, we investigated to role of catecholamines noradrenaline and dopamine on sequential information gathering. We found that blocking noradrenaline led to a decrease in information gathering, with no effect seen following dopamine blockade. Using a Bayesian computational model, we show that this noradrenaline effect is driven by an increased decision urgency, a signal that reflects an escalating subjective cost of sampling. The observation that noradrenaline modulates decision urgency suggests new avenues for treating patients that show information gathering deficits.

neuroscience

The Long and the Short of Serotonergic Stimulation: Optogenetic activation of dorsal raphe serotonergic neurons changes the learning rate for rewards

Serotonin plays an influential, but computationally obscure, modulatory role in many aspects of normal and dysfunctional learning and cognition. Here, we studied the impact of optogenetic stimulation of dorsal raphe serotonin neurons in mice performing a non-stationary, reward-driven, foraging task. We report that activation of serotonin neurons significantly boosted learning rates for choices following long inter-trial-intervals that were driven by the recent history of reinforcement.

neuroscience

Single-Trial Inhibition of Anterior Cingulate Disrupts Model-based Reinforcement Learning in a Two-step Decision Task.

The anterior cingulate cortex (ACC) is implicated in learning the value of actions, but it remains poorly understood whether and how it contributes to model-based mechanisms that use action-state predictions and afford behavioural flexibility. To isolate these mechanisms, we developed a multi-step decision task for mice in which both action-state transition probabilities and reward probabilities changed over time. Calcium imaging revealed ramps of choice-selective neuronal activity, followed by an evolving representation of the state reached and trial outcome, with different neuronal populations representing reward in different states. ACC neurons represented the current action-state transition structure, whether state transitions were expected or surprising, and the predicted state given chosen action. Optogenetic inhibition of ACC blocked the influence of action-state transitions on subsequent choice, without affecting the influence of rewards. These data support a role for ACC in model-based reinforcement learning, specifically in using action-state transitions to guide subsequent choice. HighlightsO_LIA novel two-step task disambiguates model-based and model-free RL in mice. C_LIO_LIACC represents all trial events, reward representation is contextualised by state. C_LIO_LIACC represents action-state transition structure, predicted states, and surprise. C_LIO_LIInhibiting ACC impedes action-state transitions from influencing subsequent choice. C_LI

neuroscience

Modelling avoidance in pathologically anxious humans using reinforcement-learning

Serious and debilitating symptoms of anxiety are the most common mental health problem worldwide, accounting for around 5% of all adult years lived with disability in the developed world. Avoidance behaviour -avoiding social situations for fear of embarrassment, for instance-is a core feature of such anxiety. However, as for many other psychiatric symptoms, the biological mechanisms underlying avoidance remain unclear. Reinforcement-learning models provide formal and testable characterizations of the mechanisms of decision-making; here, we examine avoidance in these terms. One hundred and one healthy and pathologically anxious individuals completed an approach-avoidance go/no-go task under stress induced by threat of unpredictable shock. We show an increased reliance in the anxious group on a parameter of our reinforcement-learning model that characterizes a prepotent (Pavlovian) bias to withhold responding in the face of negative outcomes. This was particularly the case when the anxious individuals were under stress. This formal description of avoidance within the reinforcement-learning framework provides a new means of linking clinical symptoms with biophysically plausible models of neural circuitry and, as such, takes us closer to a mechanistic understanding of pathological anxiety.

animal behavior and cognition