bioRxiv ScienceSearch

Biology subjects

Wu, C. M.

Publications and source records attributed to Wu, C. M..

6 recordsLinked to original sources

Searching for rewards like a child means less generalization and more directed exploration

How do children and adults differ in their search for rewards? We consider three different hypotheses that attribute developmental differences to either childrens increased random sampling, more directed exploration towards uncertain options, or narrower generalization. Using a search task in which noisy rewards are spatially correlated on a grid, we compare 55 younger children (age 7-8), 55 older children (age 9-11), and 50 adults (age 19-55) in their ability to successfully generalize about unobserved outcomes and balance the exploration-exploitation dilemma. Our results show that children explore more eagerly than adults, but obtain lower rewards. Building a predictive model of search to disentangle the unique contributions of the three hypotheses of developmental differences, we find robust and recoverable parameter estimates indicating that children generalize less and rely on directed exploration more than adults. We do not, however, find reliable differences in terms of random sampling.

animal behavior and cognition

Connecting conceptual and spatial search via a model of generalization

The idea of a \"cognitive map\" was originally developed to explain planning and generalization in spatial domains through a representation of inferred relationships between experiences. Recently, new research has suggested similar principles may also govern the representation of more abstract, conceptual knowledge in the brain. We test whether the search for rewards in conceptual spaces follows similar computational principles as in spatial environments. Using a within-subject design, participants searched for both spatially and conceptually correlated rewards in multi-armed bandit tasks. We use a Gaussian Process model combining generalization with an optimistic sampling strategy to capture human search decisions and judgments in both domains, and to simulate human-level performance when specified with participant parameter estimates. In line with the notion of a domain-general generalization mechanism, parameter estimates correlate across spatial and conceptual search, yet some differences also emerged, with participants generalizing less and exploiting more in the conceptual domain.

animal behavior and cognition

Sharing is not erring: Pseudo-reciprocity in collective search

Information sharing in competitive environments may seem counterintuitive, yet it is widely observed in humans and other animals. For instance, the open-source software movement has led to new and valuable technologies being released publicly to facilitate broader collaboration and further innovation. What drives this behavior and under which conditions can it be beneficial for an individual? Using simulations in both static and dynamic environments, we show that sharing information can lead to individual benefits through the mechanisms of pseudo-reciprocity, whereby shared information leads to by-product benefits for an individual without the need for explicit reciprocation. Crucially, imitation with a certain level of innovation is required to avoid a tragedy of the commons, while the mechanism of a local visibility radius allows for the coordination of self-organizing collectives of agents. When these two mechanisms are present, we find robust evidence for the benefits of sharing--even when others do not reciprocate.

animal behavior and cognition

Generalization and search in risky environments

How do people pursue rewards in risky environments, where some outcomes should be avoided at all costs? We investigate how participant search for spatially correlated rewards in scenarios where one must avoid sampling rewards below a given threshold. This requires not only the balancing of exploration and exploitation, but also reasoning about how to avoid potentially risky areas of the search space. Within risky versions of the spatially correlated multi-armed bandit task, we show that participants behavior is aligned well with a Gaussian process function learning algorithm, which chooses points based on a safe optimization routine. Moreover, using leave-one-block-out cross-validation, we find that participants adapt their sampling behavior to the riskiness of the task, although the underlying function learning mechanism remains relatively unchanged. These results show that participants can adapt their search behavior to the adversity of the environment and enrich our understanding of adaptive behavior in the face of risk and uncertainty.

animal behavior and cognition

Exploration and generalization in vast spaces

From foraging for food to learning complex games, many aspects of human behaviour can be framed as a search problem with a vast space of possible actions. Under finite search horizons, optimal solutions are generally unobtainable. Yet how do humans navigate vast problem spaces, which require intelligent exploration of unobserved actions? Using a variety of bandit tasks with up to 121 arms, we study how humans search for rewards under limited search horizons, where the spatial correlation of rewards (in both generated and natural environments) provides traction for generalization. Across a variety of diifferent probabilistic and heuristic models, we find evidence that Gaussian Process function learning--combined with an optimistic Upper Confidence Bound sampling strategy--provides a robust account of how people use generalization to guide search. Our modelling results and parameter estimates are recoverable, and can be used to simulate human-like performance, providing insights about human behaviour in complex environments.

animal behavior and cognition

Mapping the unknown: The spatially correlated multi-armed bandit

We introduce the spatially correlated multi-armed bandit as a task coupling function learning with the exploration-exploitation trade-off. Participants interacted with bi-variate reward functions on a two-dimensional grid, with the goal of either gaining the largest average score or finding the largest payoff. By providing an opportunity to learn the underlying reward function through spatial correlations, we model to what extent people form beliefs about unexplored payoffs and how that guides search behavior. Participants adapted to assigned payoff conditions, performed better in smooth than in rough environments, and--surprisingly--sometimes performed equally well in short as in long search horizons. Our modeling results indicate a preference for local search options, which when accounted for, still suggests participants were best-described as forming local inferences about unexplored regions, combined with a search strategy that directly traded off between exploiting high expected rewards and exploring to reduce uncertainty about the spatial structure of rewards.

animal behavior and cognition