bioRxiv Science⌕ Search

Biology subjects

Azadpour, M.

Publications and source records attributed to Azadpour, M..

3 recordsLinked to original sources

Employing Deep Learning Model to Evaluate Speech Information in Vocoder Simulations of Auditory Implants

Vocoder simulations have played a crucial role in the development of sound coding and speech processing techniques for auditory implant devices. Vocoders have been extensively used to model the effects of implant signal processing as well as individual anatomy and physiology on speech perception of implant users. Traditionally, such simulations have been conducted on human subjects, which can be time-consuming and costly. In addition, perception of vocoded speech varies significantly across individual subjects, and can be significantly affected by small amounts of familiarization or exposure to vocoded sounds. In this study, we propose a novel method that differs from traditional vocoder studies. Rather than using actual human participants, we use a speech recognition model to examine the influence of vocoder-simulated cochlear implant processing on speech perception. We used the OpenAI Whisper, a recently developed advanced open-source deep learning speech recognition model. The Whisper models performance was evaluated on vocoded words and sentences in both quiet and noisy conditions with respect to several vocoder parameters such as number of spectral bands, input frequency range, envelope cut-off frequency, envelope dynamic range, and number of discriminable envelope steps. Our results indicate that the Whisper model exhibited human-like robustness to vocoder simulations, with performance closely mirroring that of human subjects in response to modifications in vocoder parameters. Furthermore, this proposed method has the advantage of being far less expensive and quicker than traditional human studies, while also being free from inter-individual variability in learning abilities, cognitive factors, and attentional states. Our study demonstrates the potential of employing advanced deep learning models of speech recognition in auditory prosthesis research.

bioengineering↗

Capabilities of the CCi-Mobile Cochlear Implant Research Platform for Real-Time Sound Coding

One important obstacle to optimizing fitting and sound coding for auditory implants is lack of flexible, powerful and portable platforms that can be used in real-world listening environments by implanted patients. The clinical processors and the typically available research tools either do not have sufficient computational power and flexibility or are not portable. In response to this need, the Center for Robust Speech Systems (CRSS) at the University of Texas at Dallas has developed CCI-Mobile, in collaboration with the Laboratory for Translational Auditory Research at New York University School of Medicine and the Binaural Hearing and Speech Laboratory at the University of Wisconsin-Madison. The CCI-Mobile platform provides unique flexibility to implement a variety of real-time sound coding algorithms in real-world environments, including algorithms that require synchronized binaural stimulation. In this paper, we will describe the overall architecture of the CCI-Mobile platform and provide practical considerations for designing real-time sound coding algorithms with this platform. CCI-Mobile is under development and future generations may provide further functionality, beyond what is described in this paper.

bioengineering↗

Reducing interaural tonotopic mismatch preserves binaural unmasking in cochlear implant simulations of single-sided deafness

Binaural unmasking, a key feature of normal binaural hearing, refers to the improved intelligibility of masked speech by adding masking noise that facilities perceived spatial separation of target and masker. A question particularly relevant for cochlear implant users with single-sided deafness (SSD-CI) is whether binaural unmasking can still be achieved if the additional masking is distorted. Adding the CI restores some aspects of binaural hearing to these listeners, although binaural unmasking remains limited. Notably, these listeners may experience a mismatch between the frequency information perceived through the CI and that perceived by their normal hearing ear. Employing acoustic simulations of SSD-CI with normal hearing listeners, the present study confirms a previous simulation study that binaural unmasking is severely limited when interaural frequency mismatch between the input frequency range and simulated place of stimulation exceeds 1-2 mm. The present study also shows that binaural unmasking is largely retained when the input frequency range is adjusted to match simulated place of stimulation, even at the expense of removing low-frequency information. This result bears implication for the mechanisms driving the type of binaural unmasking of the present study, as well as for mapping the frequency range of the CI speech processor in SSD-CI users.

neuroscience↗