bioRxiv · 10.1101/2020.05.12.090134
A normative account of confirmatory biases during reinforcement learning
Abstract
Reinforcement learning involves updating estimates of the value of states and actions on the basis of experience. Previous work has shown that in humans, reinforcement learning exhibits a confirmatory bias: when updating the value of a chosen option, estimates are revised more radically following positive than negative reward prediction errors, but the converse is observed when updating the unchosen option value estimate. Here, we simulate performance on a multi-arm bandit task to examine the consequences of a confirmatory bias for reward harvesting. We report a paradoxical finding: that confirmatory biases allow the agent to maximise reward relative to an unbiased updating rule. This principle holds over a wide range of experimental settings and is most influential when decisions are corrupted by noise. We show that this occurs because on average, confirmatory biases lead to overestimating the value of more valuable bandits, and underestimating the value of less valuable bandits, rendering decisions overall more robust in the face of noise. Our results show how apparently suboptimal learning policies can in fact be reward-maximising if decisions are made with finite computational precision.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lefebvre, G., Summerfield, C., Bogacz, R.. 2020-05-14. A normative account of confirmatory biases during reinforcement learning. https://doi.org/10.1101/2020.05.12.090134
Cite the original work for its findings. Save a collection to share your selection of sources.