bioRxiv Science⌕ Search

Biology subjects

Lado, A.

Publications and source records attributed to Lado, A..

2 recordsLinked to original sources

The Noisy Encoding of Disparity Model Predicts Perception of the McGurk Effect in Native Japanese Speakers

The McGurk effect is an illusion that demonstrates the influence of information from the face of the talker on the perception of auditory speech. The diversity of human languages has prompted many intercultural studies of the effect, including in native Japanese speakers. Studies of large samples of native English speakers have shown that the McGurk effect is characterized by high variability, both in the susceptibility of different individuals to the illusion and in the frequency with which different experimental stimuli induce the illusion. The noisy encoding of disparity (NED) model of the McGurk effect uses Bayesian principles to account for this variability by separately estimating the susceptibility and sensory noise for each individual and the strength of each stimulus. To test whether the NED model could account for McGurk perception in a non-Western culture, we applied it to data collected from 80 native Japanese-speaking participants. Fifteen different McGurk stimuli were presented, along with audiovisual congruent stimuli. The McGurk effect was highly variable across stimuli and participants, with the percentage of illusory fusion responses ranging from 3% to 78% across stimuli and from 0% to 91% across participants. Despite this variability, the NED model accurately predicted perception, predicting fusion rates for individual stimuli with 2.1% error and for individual participants with 2.4% error. Stimuli containing the unvoiced pa/ka pairing evoked more fusion responses than the voiced ba/ga pairing. Model estimates of sensory noise was correlated with participant age, with greater sensory noise in older participants. The NED model of the McGurk effect offers a principled way to account for individual and stimulus differences when examining the McGurk effect within and across cultures.

neuroscience↗

The Effect on Speech-in-Noise Perception of Real Faces and Synthetic Faces Generated with either Deep Neural Networks or the Facial Action Coding System

The prevalence of synthetic talking faces in both commercial and academic environments is increasing as the technology to generate them grows more powerful and available. While it has long been known that seeing the face of the talker improves human perception of speech-in-noise, recent studies have shown that synthetic talking faces generated by deep neural networks (DNNs) are also able to improve human perception of speech-in-noise. However, in previous studies the benefit provided by DNN synthetic faces was only about half that of real human talkers. We sought to determine whether synthetic talking faces generated by an alternative method would provide a greater perceptual benefit. The facial action coding system (FACS) is a comprehensive system for measuring visually discernible facial movements. Because the action units that comprise FACS are linked to specific muscle groups, synthetic talking faces generated by FACS might have greater verisimilitude than DNN synthetic faces which do not reference an explicit model of the facial musculature. We tested the ability of human observers to identity speech-in-noise accompanied by a blank screen; the real face of the talker; and synthetic talking face generated either by DNN or FACS. We replicated previous findings of a large benefit for seeing the face of a real talker for speech-in-noise perception and a smaller benefit for DNN synthetic faces. FACS faces also improved perception, but only to the same degree as DNN faces. Analysis at the phoneme level showed that the performance of DNN and FACS faces was particularly poor for phonemes that involve interactions between the teeth and lips, such as /f/, /v/, and /th/. Inspection of single video frames revealed that the characteristic visual features for these phonemes were weak or absent in synthetic faces. Modeling the real vs. synthetic difference showed that increasing the realism of a few phonemes could substantially increase the overall perceptual benefit of synthetic faces, providing a roadmap for improving communication in this rapidly developing domain.

neuroscience↗