bioRxiv ScienceSearch

Biology subjects

Palminteri, S.

Publications and source records attributed to Palminteri, S..

8 recordsLinked to original sources

Computational noise in reward-guided learning drives behavioral variability in volatile environments

When learning the value of actions in volatile environments, humans often make seemingly irrational decisions which fail to maximize expected value. We reasoned that these non-greedy decisions, instead of reflecting information seeking during choice, may be caused by computational noise in the learning of action values. Here, using reinforcement learning (RL) models of behavior and multimodal neurophysiological data, we show that the majority of non-greedy decisions stems from this learning noise. The trial-to-trial variability of sequential learning steps and their impact on behavior could be predicted both by BOLD responses to obtained rewards in the dorsal anterior cingulate cortex (dACC) and by phasic pupillary dilation - suggestive of neuromodulatory fluctuations driven by the locus coeruleus-norepinephrine (LC-NE) system. Together, these findings indicate that most of behavioral variability, rather than reflecting human exploration, is due to the limited computational precision of reward-guided learning.

neuroscience

How the level of reward awareness changes the computational and electrophysiological signatures of reinforcement learning

The extent to which subjective awareness influences reward processing, and thereby affects future decisions is currently largely unknown. In the present report, we investigated this question in a reinforcement-learning framework, combining perceptual masking, computational modeling and electroencephalographic recordings (human male and female participants). Our results indicate that degrading the visibility of the reward decreased -without completely obliterating- the ability of participants to learn from outcomes, but concurrently increased their tendency to repeat previous choices. We dissociated electrophysiological signatures evoked by the reward-based learning processes from those elicited by the reward-independent repetition of previous choices and showed that these neural activities were significantly modulated by reward visibility. Overall, this report sheds new light on the neural computations underlying reward-based learning and decision-making and highlights that awareness is beneficial for the trial-by-trial adjustment of decision-making strategies.\n\nSignificance statementThe notion of reward is strongly associated with subjective evaluation, related to conscious processes such as \"pleasure\", \"liking\" and \"wanting\". Here we show that degrading reward visibility in a reinforcement learning task decreases -without completely obliterating- the ability of participants to learn from outcomes, but concurrently increases subjects tendency to repeat previous choices. Electrophysiological recordings, in combination with computational modelling, show that neural activities were significantly modulated by reward visibility. Overall, we dissociate different neural computations underlying reward-based learning and decision-making, which highlights a beneficial role of reward awareness in adjusting decision-making strategies.

neuroscience

Trading off the cost of conflict against expected rewards

Value-based decision-making involves trading off the cost associated with an action against its expected reward. Research has shown that both physical and mental effort constitute such subjective costs, biasing choices away from effortful actions, and discounting the value of obtained rewards. Facing conflicts between competing action alternatives is considered aversive, as recruiting cognitive control to overcome conflict is effortful. Yet, it remains unclear whether conflict is also perceived as a cost in value-based decisions. The present study investigated this question by embedding irrelevant distractors (flanker arrows) within a reversal-learning task, with intermixed free and instructed trials. Results showed that participants learned to adapt their choices to maximize rewards, but were nevertheless biased to follow the suggestions of irrelevant distractors. Thus, the perceived cost of being in conflict with an external suggestion could sometimes trump internal value representations. By adapting computational models of reinforcement learning, we assessed the influence of conflict at both the decision and learning stages. Modelling the decision showed that conflict was avoided when evidence for either action alternative was weak, demonstrating that the cost of conflict was traded off against expected rewards. During the learning phase, we found that learning rates were reduced in instructed, relative to free, choices. Learning rates were further reduced by conflict between an instruction and subjective action values, whereas learning was not robustly influenced by conflict between ones actions and external distractors. Our results show that the subjective cost of conflict factors into value-based decision-making, and highlights that different types of conflict may have different effects on learning about action outcomes.

neuroscience

Social information impairs reward learning in depressive subjects: behavioral and computational characterization

Depression is characterized by a marked decrease in social interactions and blunted sensitivity to rewards. Surprisingly, despite the importance of social deficits in depression, non-social aspects have been disproportionally investigated. As a consequence, the cognitive mechanisms underlying atypical decision-making in social contexts in depression are poorly understood. In the present study, we investigate whether deficits in reward processing interact with the social context and how this interaction is affected by self-reported depression and anxiety symptoms. Two cohorts of subjects (discovery and replication sample: N = 50 each) took part in a task involving reward learning in a social context with different levels of social information (absent, partial and complete). Behavioral analyses revealed a specific detrimental effect of depressive symptoms - but not anxiety - on behavioral performance in the presence of social information, i.e. when participants were informed about the choices of another player. Model-based analyses further characterized the computational nature of this deficit as a negative audience effect, rather than a deficit in the way others choices and rewards are integrated in decision making. To conclude, our results shed light on the cognitive and computational mechanisms underlying the interaction between social cognition, reward learning and decision-making in depressive disorders.

neuroscience

Contextual influence on confidence judgments in human reinforcement learning.

The ability to correctly estimate the probability of ones choices being correct is fundamental to optimally re-evaluate previous choices or to arbitrate between different decision strategies. Experimental evidence nonetheless suggests that this metacognitive process -referred to as a confidence judgment-is susceptible to numerous biases. We investigate the effect of outcome valence (gains or losses) on confidence while participants learned stimulus-outcome associations by trial-and-error. In two experiments, we demonstrate that participants are more confident in their choices when learning to seek gains compared to avoiding losses. Importantly, these differences in confidence were observed despite objectively equal choice difficulty and similar observed performance between those two contexts. Using computational modelling, we show that this bias is driven by the context-value, a dynamically updated estimate of the average expected-value of choice options that has previously been demonstrated to be necessary to explain equal performance in the gain and loss domain. The biasing effect of context-value on confidence, also recently observed in the context of incentivized perceptual decision-making, is therefore domain-general, with likely important functional consequences.

animal behavior and cognition

Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences

In economics and in perceptual decision-making contextual effects are well documented, where decision weights are adjusted as a function of the distribution of stimuli. Yet, in reinforcement learning literature whether and how contextual information pertaining to decision states is integrated in learning algorithms has received comparably little attention. Here, in an attempt to fill this gap, we investigated reinforcement learning behavior and its computational substrates in a task where we orthogonally manipulated both outcome valence and magnitude, resulting in systematic variations in state-values. Over two experiments, model comparison indicated that subjects behavior is best accounted for by an algorithm which includes both reference point-dependence and range-adaptation - two crucial features of state-dependent valuation. In addition, we found state-dependent outcome valuation to progressively emerge over time, to be favored by increasing outcome information and to be correlated with explicit understanding of the task structure. Finally, our data clearly show that, while being locally adaptive (for instance in negative valence and small magnitude contexts), state-dependent valuation comes at the cost of seemingly irrational choices, when options are extrapolated out from their original contexts.

animal behavior and cognition

Neural Computations Underpinning The Strategic Management Of Influence In Advice Giving

Research on social influence has mainly focused on the target of influence (e.g. consumer, voter), while the cognitive and neurobiological underpinnings of the source of the influence (e.g. marketer, politician) remain unexplored. Here we show that advisers managed their influence over their client strategically, by modulating the confidence of their advice, depending on the interaction between advisers level of influence on the client (i.e., being ignored or chosen by the client), and relative merit (i.e., their accuracy in comparison with a rival). Functional magnetic resonance imaging showed that these sources of social information were tracked in distinct regions of the brains social and valuation systems: relative merit in the medial-prefrontal cortex, and selection by client in the temporo-parietal junction. In addition, increased functional connectivity between these two regions during positive merit trials predicted their effect on behaviour. Both sources of social information modulated the activity in the ventral striatum. These results further our understanding of human interactions and provide a framework for investigating the neurobiology of how we try to influence others.

neuroscience

Confirmation bias in human reinforcement learning: evidence from counterfactual feedback processing

Previous studies suggest that factual learning, that is learning from obtained outcomes, is biased, such that participants preferentially take into account positive, as compared to negative, prediction errors. However, whether or not the prediction error valence also affects counterfactual learning, that is, learning from forgone outcomes, is unknown. To address this question, we analysed the performance of two cohorts of participants on reinforcement learning tasks using a computational model that was adapted to test if prediction error valance influences learning. Concerning factual learning, we replicated previous findings of a valence-induced bias, whereby participants learned preferentially from positive, relative to negative, prediction errors. In contrast, for counterfactual learning, we found the opposite valence-induced bias: negative prediction errors were preferentially taken into account relative to positive ones. When considering valence-induced bias in the context of both factual and counterfactual learning, it appears that people tend to preferentially take into account information that confirms their current choice. By documenting these valence-induced learning biases, our findings demonstrate the presence of a confirmation bias in human reinforcement learning.

animal behavior and cognition