bioRxiv Science⌕ Search

Biology subjects

McNamee, D.

Publications and source records attributed to McNamee, D..

4 recordsLinked to original sources

Expert Navigators Deploy Rational Hierarchical Priorization Over Predictive Maps For Large-Scale Real-World Planning

Efficient planning is a distinctive hallmark of intelligence in humans, who routinely make rapid inferences over complex world contexts. However, studies investigating how humans accomplish this tend to focus on naive participants engaged in simplistic tasks with small state-spaces, which do not reflect the intricacy, ecological validity, and human specialisation in real-world planning. In this study, we examine the street-by-street route planning of London taxi drivers navigating across more than 26,000 streets in London (UK). We explore how planning unfolded dynamically over different phases of journey construction and identify theoretic principles by which these expert human planners rationally precache decisions at prioritised environment states in an early phase of the planning process. In particular, we find that measures of path complexity predict human mental sampling prioritisation dynamics independent of alternative measures derived from the real spatial context being navigated. Our data provide real-world evidence for complexity-driven remote state access within internal models and precaching during human expert route planning in very large structured spaces. Significance statementHumans can plan efficiently in incredibly complex situations. Existing work has looked at naive participants in simple tasks, which might not be representative of how experts plan in the real world. Here, we study the real-world planning process of London taxi drivers - famous for their expert knowledge of London. By analyzing their response times as a proxy for thinking times, we reveal that at an early stage in their thought process, they store decisions at key street junctions to keep them in mind for later planning. Using computational modeling, we show that taxi drivers prioritize inference at street junctions according to normative metrics measuring how critical a particular decision is for reducing the complexity of planning across the entire city.

neuroscience↗

London taxi drivers exploit neighbourhood boundaries for hierarchical route planning

Humans show an impressive ability to plan over complex situations and environments. A classic approach to explaining such planning has been tree-search algorithms which search through alternative state sequences for the most efficient path through states. However, this approach fails when the number of states is large due to the time to compute all possible sequences. Hierarchical route planning has been proposed as an alternative, offering a computationally efficient mechanism in which the representation of the environment is segregated into clusters. Current evidence for hierarchical planning comes from experimentally created environments which have clearly defined boundaries and far fewer states than the real-world. To test for real-world hierarchical planning we exploited the capacity of London licensed taxi drivers to use their memory to construct a street by street plan across London, UK (>26,000 streets). The time to recall each successive street name was treated as the response time, with a rapid average of 1.8 seconds between each street. In support of hierarchical planning we find that the clustered structure of Londons regions impacts the response times, with minimal impact of the distance across the street network (as would be predicted by tree-search). We also find that changing direction during the plan (e.g. turning left or right) is associated with delayed response times. Thus, our results provide real-world evidence for how humans structure planning over a very large number of states, and give a measure of human expertise in planning.

neuroscience↗

Dopamine neurons encode a multidimensional probabilistic map of future reward

Learning to predict rewards is a fundamental driver of adaptive behavior. Midbrain dopamine neurons (DANs) play a key role in such learning by signaling reward prediction errors (RPEs) that teach recipient circuits about expected rewards given current circumstances and actions. However, the algorithm that DANs are thought to provide a substrate for, temporal difference (TD) reinforcement learning (RL), learns the mean of temporally discounted expected future rewards, discarding useful information concerning experienced distributions of reward amounts and delays. Here we present time-magnitude RL (TMRL), a multidimensional variant of distributional reinforcement learning that learns the joint distribution of future rewards over time and magnitude using an efficient code that adapts to environmental statistics. In addition, we discovered signatures of TMRL-like computations in the activity of optogenetically identified DANs in mice during a classical conditioning task. Specifically, we found significant diversity in both temporal discounting and tuning for the magnitude of rewards across DANs, features that allow the computation of a two dimensional, probabilistic map of future rewards from just 450ms of neural activity recorded from a population of DANs in response to a reward-predictive cue. In addition, reward time predictions derived from this population code correlated with the timing of anticipatory behavior, suggesting the information is used to guide decisions regarding when to act. Finally, by simulating behavior in a foraging environment, we highlight benefits of access to a joint probability distribution of reward over time and magnitude in the face of dynamic reward landscapes and internal physiological need states. These findings demonstrate surprisingly rich probabilistic reward information that is learned and communicated to DANs, and suggest a simple, local-in-time extension of TD learning algorithms that explains how such information may be acquired and computed.

animal behavior and cognition↗

Distinct replay signatures for planning and memory maintenance

Theories of neural replay propose that it supports a range of functions, most prominently planning and memory consolidation. Here, we test the hypothesis that distinct signatures of replay in the same task are related to model-based decisionmaking ( planning) and memory preservation. We designed a reward learning task wherein participants utilized structure knowledge for model-based evaluation, while at the same time had to maintain knowledge of two independent and randomly alternating task environments. Using magnetoencephalography (MEG) and multivariate analysis, we first identified temporally compressed sequential reactivation, or replay, both prior to choice and following reward feedback. Before choice, prospective replay strength was enhanced for the current task-relevant environment when a model-based planning strategy was beneficial. Following reward receipt, and consistent with a memory preservation role, replay for the alternative distal task environment was enhanced as a function of decreasing recency of experience with that environment. Critically, these planning and memory preservation relationships were selective to pre-choice and post-feedback periods. Our results provide new support for key theoretical proposals regarding the functional role of replay and demonstrate that the relative strength of planning and memory-related signals are modulated by on-going computational and task demands. Significance statementThe sequential neural reactivation of prior experience, known as replay, is considered to be an important mechanism for both future planning and preserving memories of the past. Whether, and how, replay supports both of these functions remains unknown. Here, in humans, we found that prior to a choice, rapid replay of potential future paths was enhanced when planning was more beneficial. By contrast, after choice feedback, when no future actions are imminent, we found evidence for a memory preservation signal evident in enhanced replay of paths that had been visited less in the recent past. The results demonstrate that distinct replay signatures, expressed at different times, relate to two dissociable cognitive functions.

neuroscience↗