bioRxiv · 10.1101/2023.09.30.560270
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Abstract
This paper introduces WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for human and animal Voice Activity Detection (VAD). Contrary to traditional methods that detect human voice or animal vocalizations from a short audio frame and rely on careful threshold selection, WhisperSeg processes entire spectrograms of long audio and generates plain text representations of onset, offset, and type of voice activity. Processing a longer audio context with a larger network greatly improves detection accuracy from few labeled examples. We further demonstrate a positive transfer of detection performance to new animal species, making our approach viable in the data-scarce multi-species setting.1
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gu, N., Lee, K., Basha, M., Ram, S. K., You, G., Hahnloser, R.. 2023-10-02. Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection. https://doi.org/10.1101/2023.09.30.560270
Cite the original work for its findings. Save a collection to share your selection of sources.