bioRxiv Science⌕ Search

Biology subjects

Tchernichovski, O.

Publications and source records attributed to Tchernichovski, O..

1 recordsLinked to original sources

Learning the action inventory of a complex skill via an intrinsic reward

Reinforcement learning (RL) is thought to underlie the acquisition of vocal skills like birdsong and speech, where sounding like ones "tutor" is rewarding. Yet, we find that the standard actor-critic RL model of birdsong learning fails to account for juvenile zebra finches efficient learning of an inventory of multiple syllables. But when we replace a single actor with multiple independent actors that jointly maximize a common intrinsic reward, then birds empirical learning trajectories are accurately reproduced. Importantly, the influence of each actor (syllable) on the magnitude of global reward is competitively determined by its acoustic similarity to target syllables. This leads to each actor matching the target it is closest to, and occasionally, to the competitive exclusion of an actor from the learning process (i.e., the learned song). We propose that a competitive-cooperative multi-actor (MARL) algorithm is key for the efficient learning of the action inventory of a complex skill.

animal behavior and cognition↗