bioRxiv ScienceSearch

Biology subjects

O'Sullivan, A. E.

Publications and source records attributed to O'Sullivan, A. E..

3 recordsLinked to original sources

The integration of continuous audio and visual speech in a cocktail-party environment depends on attention

In noisy, complex environments, our ability to understand audio speech benefits greatly from seeing the speakers face. This is attributed to the brains ability to integrate audio and visual information, a process known as multisensory integration. In addition, selective attention to speech in complex environments plays an enormous role in what we understand, the so-called cocktail-party phenomenon. But how attention and multisensory integration interact remains incompletely understood. While considerable progress has been made on this issue using simple, and often illusory (e.g., McGurk) stimuli, relatively little is known about how attention and multisensory integration interact in the case of natural, continuous speech. Here, we addressed this issue by analyzing EEG data recorded from subjects who undertook a multisensory cocktail-party attention task using natural speech. To assess multisensory integration, we modeled the EEG responses to the speech in two ways. The first assumed that audiovisual speech processing is simply a linear combination of audio speech processing and visual speech processing (i.e., an A+V model), while the second allows for the possibility of audiovisual interactions (i.e., an AV model). Applying these models to the data revealed that EEG responses to attended audiovisual speech were better explained by an AV model than an A+V model, providing evidence for multisensory integration. In contrast, unattended audiovisual speech responses were best captured using an A+V model, suggesting that multisensory integration is suppressed for unattended speech. Follow up analyses revealed some limited evidence for early multisensory integration of unattended AV speech, with no integration occurring at later levels of processing. We take these findings as evidence that the integration of natural audio and visual speech occurs at multiple levels of processing in the brain, each of which can be differentially affected by attention.

neuroscience

A linguistic representation in the visual system underlies successful lipreading

There is considerable debate over how visual speech is processed in the absence of sound and whether neural activity supporting lipreading occurs in visual brain areas. Surprisingly, much of this ambiguity stems from a lack of behaviorally grounded neurophysiological findings. To address this, we conducted an experiment in which human observers rehearsed audiovisual speech for the purpose of lipreading silent versions during testing. Using a combination of computational modeling, electroencephalography, and simultaneously recorded behavior, we show that the visual system produces its own specialized representation of speech that is 1) well-described by categorical linguistic units ("visemes") 2) dissociable from lip movements, and 3) predictive of lipreading ability. These findings contradict a long-held view that visual speech processing co-opts auditory cortex after early visual processing stages. Consistent with hierarchical accounts of visual and audiovisual speech perception, our findings show that visual cortex performs at least a basic level of linguistic processing.

neuroscience

Neurophysiological indices of audiovisual speech integration are enhanced at the phonetic level for speech in noise

Seeing a speakers face benefits speech comprehension, especially in challenging listening conditions. This perceptual benefit is thought to stem from the neural integration of visual and auditory speech at multiple stages of processing, whereby movement of a speakers face provides temporal cues to auditory cortex, and articulatory information from the speakers mouth can aid recognizing specific linguistic units (e.g., phonemes, syllables). However it remains unclear how the integration of these cues varies as a function of listening conditions. Here we sought to provide insight on these questions by examining EEG responses to natural audiovisual, audio, and visual speech in quiet and in noise. Specifically, we represented our speech stimuli in terms of their spectrograms and their phonetic features, and then quantified the strength of the encoding of those features in the EEG using canonical correlation analysis. The encoding of both spectrotemporal and phonetic features was shown to be more robust in audiovisual speech responses then what would have been expected from the summation of the audio and visual speech responses, consistent with the literature on multisensory integration. Furthermore, the strength of this multisensory enhancement was more pronounced at the level of phonetic processing for speech in noise relative to speech in quiet, indicating that listeners rely more on articulatory details from visual speech in challenging listening conditions. These findings support the notion that the integration of audio and visual speech is a flexible, multistage process that adapts to optimize comprehension based on the current listening conditions. Significance StatementDuring conversation, visual cues impact our perception of speech. Integration of auditory and visual speech is thought to occur at multiple stages of speech processing and vary flexibly depending on the listening conditions. Here we examine audiovisual integration at two stages of speech processing using the speech spectrogram and a phonetic representation, and test how audiovisual integration adapts to degraded listening conditions. We find significant integration at both of these stages regardless of listening conditions, and when the speech is noisy, we find enhanced integration at the phonetic stage of processing. These findings provide support for the multistage integration framework and demonstrate its flexibility in terms of a greater reliance on visual articulatory information in challenging listening conditions.

neuroscience