bioRxiv · 10.1101/2023.09.24.559177
Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation
Abstract
To transform speech into words, the human brain must accommodate variability across utterances in intonation, speech rate, volume, accents and so on. A promising approach to explaining this process has been to model electroencephalogram (EEG) recordings of brain responses to speech. Contemporary models typically invoke speech categories (e.g. phonemes) as an intermediary representational stage between sounds and words. However, such categorical models are typically hand-crafted and therefore incomplete because they cannot speak to the neural computations that putatively underpin categorization. By providing end-to-end accounts of speech-to-language transformation, new deep-learning systems could enable more complete brain models. We here model EEG recordings of audiobook comprehension with the deep-learning system Whisper. We find that (1) Whisper provides an accurate, self-contained EEG model of speech-to-language transformation; (2) EEG modeling is more accurate when including prior speech context, which pure categorical models do not support; (3) EEG signatures of speech-to-language transformation depend on listener-attention.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anderson, A. J., Davis, C., Lalor, E. C.. 2023-09-24. Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation. https://doi.org/10.1101/2023.09.24.559177
Cite the original work for its findings. Save a collection to share your selection of sources.