bioRxiv · 10.1101/439885
Computational noise in reward-guided learning drives behavioral variability in volatile environments
Abstract
When learning the value of actions in volatile environments, humans often make seemingly irrational decisions which fail to maximize expected value. We reasoned that these non-greedy decisions, instead of reflecting information seeking during choice, may be caused by computational noise in the learning of action values. Here, using reinforcement learning (RL) models of behavior and multimodal neurophysiological data, we show that the majority of non-greedy decisions stems from this learning noise. The trial-to-trial variability of sequential learning steps and their impact on behavior could be predicted both by BOLD responses to obtained rewards in the dorsal anterior cingulate cortex (dACC) and by phasic pupillary dilation - suggestive of neuromodulatory fluctuations driven by the locus coeruleus-norepinephrine (LC-NE) system. Together, these findings indicate that most of behavioral variability, rather than reflecting human exploration, is due to the limited computational precision of reward-guided learning.
Source connections
Explore related subjects
Keep this discovery
Findling, C., Skvortsova, V., Dromnelle, R., Palminteri, S., Wyart, V.. 2018-10-11. Computational noise in reward-guided learning drives behavioral variability in volatile environments. https://doi.org/10.1101/439885
Cite the original work for its findings. Save a collection to share your selection of sources.