bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.06.15.732293

Nonlinear influence of reward volatility on arbitration between multiple learning strategies reflects cost-benefit optimization

Abstract

Action selection involves two systems: a model-free reinforcement learning strategy, which relies on experience with action-outcome pairs, and a model-based reinforcement learning strategy, which enables more flexible behavior via inference using a model of the invariant environmental structure. Although environmental change requires more flexible behavior, the ability of volatility, a higher-order statistic that captures how rapidly or frequently the environment changes, to systematically modulate these strategies remains unclear. We examined the effects of reward volatility on arbitration between model-free and model-based reinforcement learning strategies using two modified two-step decision tasks. In Experiment 1, participants performed tasks with different levels of reward volatility and time pressure. In Experiment 2, we systematically manipulated reward volatility across a broader range to assess the relationship between volatility and learning strategy. Behavioral data were analyzed using model-agnostic one-trial and multitrial back analyses, reinforcement learning simulations, and hierarchical Bayesian model fitting. Across experiments, reward volatility exerted an inverse U-shaped nonlinear effect on the arbitration between model-free and model-based reinforcement learning strategies, as the model-based learning strategy was strongly driven at intermediate levels of reward volatility. These modulation effects were observed only in individuals who had learned the transition structure in the task, whereas those who had not learned the transition structure relied on the model-free learning strategy regardless of reward volatility. Reinforcement learning simulations revealed that the relative advantage of the model-based learning strategy over the model-free learning strategy peaked at intermediate levels of reward volatility. Additionally, increased time pressure shifted behavior toward the model-free learning strategy. These results demonstrated that, humans do not always use the model-based reinforcement learning strategy in uncertain and dynamic environments, even when they are aware of the task structure, supporting cost-benefit optimization. Author SummaryThe ability to flexibly guide behavior by carefully considering future consequences is fundamental to a prominent property of human intelligence and rationality. However, what drives this deliberative system? In this study, we investigated the factors that promote deliberative versus habitual behavior using decision-making tasks with uncertain structures and changing rewards. We found that participants who spontaneously learned the hidden transition structure in the task used this knowledge to guide deliberative behavior. Conversely, participants who did not learn the structure relied primarily on habitual strategies, repeating actions that had previously been rewarded. Among participants who learned the structure, the degree of deliberative behavior changed nonlinearly with reward volatility, in which the speed at which rewards changed over time. We also observed that limiting the decision time reduced deliberative behavior and promoted habitual responding. These findings suggest that under uncertain and dynamic environments, deliberative control is adaptively regulated according to cost-benefit optimization. Our results contribute to understanding how humans flexibly adjust their behavioral control systems in response to environmental conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yamada, T., Samejima, K.. 2026-06-19. Nonlinear influence of reward volatility on arbitration between multiple learning strategies reflects cost-benefit optimization. https://doi.org/10.64898/2026.06.15.732293

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Automated pup-level analysis reveals distinct effects of prenatal CBD and THC exposure on maternal retrieval

This study examined how prenatal CBD and {Delta}9-tetrahydrocannabinol (THC) exposure affects maternal caregiving in C57BL/6J mice. We developed the Machine-Automated Scoring of the Pup Retrieval Test (MAS-PRT) to overcome limitations of manual behavioral scoring. MAS-PRT integrates Multi-Animal DeepLabCut for dam and pup tracking with Detectron2 for dynamic nest reconstruction. Validated against 170 manually annotated retrieval trials, the pipeline showed high concordance with manual measurements and enabled reproducible extraction of encounter latency, retrieval latency, and locomotor trajectories. Prenatal exposure did not impair general nest-building or overall home-cage maternal care. However, pups from both CBD and THC groups showed reduced body weight at postnatal day 5. Cox proportional hazards modeling revealed divergent effects by compound: CBD-exposed dams exhibited a weight-dependent increase in probability of encountering and retrieving lighter pups, independent of pup sex. Spatial tracking further showed that dams traversed significantly shorter total trajectories when retrieving female progeny exposed to either compound. Longitudinal trial-by-trial analysis indicated intact task acquisition in controls and CBD dams, whereas THC dams displayed a flattened learning curve driven by lower retrieval latencies on initial trials. Together, these findings indicate that prenatal cannabinoid exposure does not produce generalized disruption of maternal care but instead induces compound-specific alterations in maternal reactivity, retrieval kinematics, and learning dynamics.

animal behavior and cognition↗

A comparison of female competitive traits: Female aggression peaks at nest building but female song spans multiple contexts in a temperate songbird

Female-female competition is increasingly recognized as a key driver of female ornamentation, including birdsong, which often functions in intrasexual competition. However, the specific resources females use elaborate traits to compete for remain unclear. In addition, few studies have simultaneously investigated the use of multiple competitive traits in females, despite growing independent interest in these traits (e.g., female song and aggression). We investigated the competitive contexts of female song, aggression and calling behavior in northern house wrens (Troglodytes aedon) to determine which resources females compete for across the breeding season. We simulated conspecific territorial intrusions using female song at three breeding stages representing different contexts: arrival (mate and territory acquisition), nest building (nest site and breeding status defense), and egg laying (brood defense). We tested whether female song and physical aggression varied as reproductive resources shifted across the breeding cycle. Females were significantly more aggressive during nest building, showing 5.8 times greater odds of a higher-intensity aggressive response during nest building compared to arrival. Female song output was similar across early stages but declined during egg laying, though this was not statistically significant after correction for multiple comparisons and individuals varied substantially in overall singing propensity. Non-song vocalizations varied by call type and breeding stage. Calls associated with aggression occurred most frequently during nest building, consistent with peak physical aggression responses. Together, these results identify nest building as the stage of highest female aggression, consistent with heightened competition over nest cavities and associated breeding status in this cavity-nesting species. In contrast, female song occurred across all stages and appears to function in multiple competitive contexts. This study provides evidence for context and mode-specific female signaling in a temperate songbird and highlights that females strategically use aggression, calls, and song to mediate social conflict across breeding contexts.

animal behavior and cognition↗

Tracking human foragers and their prey reveals adaptive predator-prey dynamics

Hunting for mobile prey is thought to have played a key role in hominin evolution, by providing high-quality nutrition that supported the development of the exceptionally large human brain. However, human-prey dynamics remain poorly understood because studies have not yet tracked human foragers and their prey simultaneously. Here, we employ high resolution tracking of groups of human foragers (ice-fishers) and their prey (fish shoals) to study human-prey dynamics. Our results show that foragers adaptively combined personal and social information in deciding where to forage and for how long, closely matching the prey distribution. Prey responded dynamically to human exploitation, showing increased attraction to fishing activity, alongside decreased biting probability. Furthermore, we found that foragers adaptively relied on memory, preferentially returning to areas with high prey presence, particularly when their current return rate was low. Our results show how human foragers overcome the challenges of extracting invisible, mobile and reactive prey by tightly tuning patch-selection, patch-leaving and patch-return decisions to the distribution and behaviour of their prey.

animal behavior and cognition↗