Unifying error and reward action learning: a cerebello-basal ganglia theory
Learning depends on both reward- and error-based feedback, yet how the brain integrates these distinct signals to guide behaviour remains fundamentally unclear. Here, using a normative computational framework, we derive credit assignment rules for both reward-based learning (RBL) and error-based learning (EBL). In contrast to existing dual-policy accounts, our approach demonstrates that RBL and EBL updates can be reformulated into a shared action-gradient space that directly updates a single downstream policy. First, we map this action-gradient framework onto a systems-level account of coordinated interactions between the basal ganglia, cerebellum, and cortex. The model reproduces key behavioral features across both learning regimes, generates experimentally testable predictions, and provides a unified computational account of motor deficits observed in patients with cerebellar and basal ganglia disorders. Together, our work offers a normative, brain-wide framework for how distributed brain systems integrate reinforcement and error-driven feedback toward a common behavioral objective.