bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.11.07.622528

Trial-by-trial learning of successor representations in human behavior

Abstract

Decisions in humans and other organisms depend, in part, on learning and using models that capture the statistical structure of the world, including the long-run expected outcomes of our actions. One prominent approach to forecasting such long-run outcomes is the successor representation (SR), which predicts future states aggregated over multiple timesteps. Although much behavioral and neural evidence suggests that people and animals use such a representation, it remains unknown how they acquire it. It has frequently been assumed to be learned by temporal difference bootstrapping (SR-TD(0)), but this assumption has largely not been empirically tested or compared to alternatives including eligibility traces (SR-TD({lambda} > 0)). Here we address this gap by leveraging trial-by-trial reaction times in graph sequence learning tasks, which are favorable for studying learning dynamics because the long horizons in these studies differentiate the transient update dynamics of different learning rules. We examined the behavior of SR-TD({lambda}) on a probabilistic graph learning task alongside a number of alternatives, and found that behavior was best explained by a hybrid model which learned via SR-TD({lambda}) alongside an additional predictive model of recency. The relatively large{lambda} we estimate indicates a predominant role of eligibility trace mechanisms over the bootstrap-based chaining typically assumed. Our results provide insight into how humans learn predictive representations, and demonstrate that people simultaneously learn the SR alongside lower-order predictions. Author SummaryOur ability to plan intelligently requires predicting the state of the world multiple steps into the future. Enumerating future outcomes step-by-step, however, is slow and costly. Instead, research has shown that people rely on simplified models of the world that skip across multiple steps at once. How do we construct these simplified models? One promising idea is the successor representation (SR), which predicts future events via a simple and neurally plausible computation. The SR has been shown to explain a range of behavioral phenomena, but these studies have not identified which among many learning rules the brain uses to build the SR. Plausible mechanisms for learning associations over delays (called bootstrapping and eligibility traces) both converge to identical simplified world models, and thus existing studies on the SR, which focus on well trained behavior, are unable to distinguish between them. Here, we answer this question by examining behavior on a graph learning task, where stimulus-by-stimulus reaction times have been shown to reflect predictions over long temporal horizons. Through both model fitting and model-agnostic comparisons, we find that behavior is best explained by a learning rule heavily dependent on eligibility traces, in contrast to previous work which generally assumed an (untested) bootstrapping update rule.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kahn, A. E., Bassett, D. S., Daw, N. D.. 2024-11-07. Trial-by-trial learning of successor representations in human behavior. https://doi.org/10.1101/2024.11.07.622528

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Attention Across Scales: From Individual Variation to Social Hierarchies and Brain Networks in Semi-Free-Ranging Macaques

Attention is a fundamental brain function supporting perception, decision-making, and social behavior, and its dysfunction profoundly impairs daily life. It is both dynamic and stable, varying across observations and individuals, changing across the lifespan, and being shaped by social and environmental experience. Yet capturing this complexity remains a central challenge in neuroscience. Here, we integrated longitudinal behavioral assessments of semi-free-ranging macaques living in naturalistic social groups with resting-state fMRI. We quantified performance across days, ages, and social hierarchies and related it to intrinsic brain organization. Distinct attentional phenotypes emerged, including individuals with reduced attentional control. Performance followed an inverted-U lifespan trajectory, improving from childhood to adulthood before declining. Social status modulated attentional performance. Critically, nonlinear lifespan trajectories and associations with individual attentional differences were most clearly expressed in frontoparietal connectivity. Together, these findings reveal how sustained attention is organized across scales, providing a biological framework for its individual diversity, social modulation, and neural basis.

neuroscience↗

Decoding natural scenes from patterned optogenetic responses in mouse visual cortex

A central challenge in developing visual cortical prostheses is to determine how visual stimuli should be transformed into effective patterns of cortical stimulation. Although advances in stimulation technologies, including optogenetics, provide increasingly precise control over cortical activity, it remains unclear whether artificially evoked activity can reproduce the information content of naturally evoked visual representations. Here we establish a quantitative framework for evaluating visual encoding strategies by decoding cortical responses evoked by natural vision and patterned optogenetic stimulation. We developed a novel dual-modal paradigm in awake mice to bridge the gap between endogenous photostimulation and artificial network driving. By co-expressing the high-performance calcium indicator GCaMP6s and the red-shifted, ultra-sensitive opsin rsChRmine-oScarlet in the primary visual cortex (V1), we successfully translated dynamic natural movie frames into patterned, spatiotemporal optogenetic stimulation. Quantitative comparisons of macro-scale dynamics demonstrated that this patterned optogenetic injection evokes cortical states highly comparable and representationally aligned with those driven by actual visual photostimulation. To systematically evaluate the fidelity of these responses, we developed STAR, a deep learning model featuring spatial and temporal attention mechanisms, and successfully reconstructed the frames of natural movies from V1 signals under both experimental modalities. Collectively, our results demonstrate that complex sensory information can be both naturally encoded and synthetically injected into V1 circuits with high decoding fidelity. This work provides an empirical and computational proof-of-concept for intelligent, closed-loop biomimetic encoders, establishing a robust framework for next-generation cortical visual neuroprostheses and bidirectional brain-machine interfaces.

neuroscience↗

Why Is Spontaneous Blink Timing Informative? An Adaptive Scheduling Perspective

Spontaneous eye blinks have long been linked to cognitive processing, yet how task demands shape blink timing and its relationship to behavioral performance remains unclear. We examined spontaneous blink behavior in 576 adults performing two variants of the Continuous Performance Task (CPT). Blink occurrence and timing were most strongly modulated by the experimental condition in the more demanding CPT-AX task, whereas their association with response time was stronger in the CPT-X task, where more consistent blink timing predicted faster responses. This dissociation suggests that task structure changes not only blink behavior but also the behavioral relevance of blink timing. These findings are consistent with an adaptive scheduling account of spontaneous blinking and provide a conceptual framework for understanding when and why blink timing contains chronometric information about ongoing cognition.

neuroscience↗