Reinforcement learning discovers new mechanisms of reentry in excitable media
The transition from transient excitation to sustained reentry is a fundamental problem in the physics of excitable media. In cardiac tissue, reentry underlies many life-threatening cardiac arrhythmias, yet the pathway to initiation of reentry remains incompletely understood. Here, we formulate reentry initiation as a reinforcement-learning problem in which an agent applies sequences of spatial stimulation patterns while being rewarded for sustained activity and penalized according to the number of stimuli applied. Using cellular automata in one-, two-, and three-dimensional geometries, the agent discovered several mechanisms for generating unidirectional propagation and reentry. These included a previously described mechanism combining superthreshold and subthreshold stimulation, as well as two new mechanisms based entirely on subthreshold stimuli: a sequential mechanism involving stimuli delivered at different locations and times, and a spatial mechanism in which several individually subthreshold sites collectively initiated reentry. In geometries containing boundaries and branches, the learned protocols additionally exploited structural source-sink asymmetries. Optogenetic experiments in cardiac monolayers further demonstrated reproducible induction of unidirectional propagation using the learned spatial subthreshold patterns, while whole-heart experiments provided preliminary evidence that such patterns can shape early propagation in intact tissue. More broadly, we show that reinforcement learning provides a general framework for discovering mechanisms of reentry in arbitrary geometries and generating testable hypotheses about reentry initiation in excitable systems.