bioRxiv Science⌕ Search

Biology subjects

Setia, T. M.

Publications and source records attributed to Setia, T. M..

2 recordsLinked to original sources

Automated detection of Bornean white-bearded gibbon (Hylobates albibarbis) vocalisations using an open-source framework for deep learning

Passive acoustic monitoring is a promising tool for monitoring at-risk populations of vocal species, yet extracting relevant information from large acoustic datasets can be time-consuming, creating a bottleneck at the point of analysis. To address this, we adapted an open-source framework for deep learning in bioacoustics to automatically detect Bornean white-bearded gibbon (Hylobates albibarbis) "great call" vocalisations in a long-term acoustic dataset from a rainforest location in Borneo. We describe the steps involved in developing this solution, including collecting audio recordings, developing training and testing datasets, training neural network models, and evaluating model performance. Our best model performed at a satisfactory level (F score = 0.87), identifying 98% of the highest-quality calls from 90 hours of manually-annotated audio recordings and greatly reduced analysis times when compared to a human observer. We found no significant difference in the temporal distribution of great call detections between the manual annotations and the models output. Future work should seek to apply our model to long-term acoustic datasets to understand spatiotemporal variations in H. albibarbis calling activity. Overall, we present a roadmap for applying deep learning to identify the vocalisations of species of interest which can be adapted for monitoring other endangered vocalising species.

ecology↗

Vocal complexity in the long calls of Bornean orangutans

Vocal complexity is central to many evolutionary hypotheses about animal communication. Yet, quantifying and comparing complexity remains a challenge, particularly when vocal types are highly graded. Male Bornean orangutans (Pongo pygmaeus wurmbii) produce complex and variable "long call" vocalizations comprising multiple sound types that vary within and among individuals. Previous studies described six distinct call (or pulse) types within these complex vocalizations, but none quantified their discreteness or the ability of human observers to reliably classify them. We studied the long calls of 13 individuals to: 1) evaluate and quantify the reliability of audio-visual classification by three well-trained observers, 2) distinguish among call types using supervised classification and unsupervised clustering, and 3) compare the performance of different feature sets. Using 46 acoustic features, we applied machine learning (i.e., support vector machines, affinity propagation, and fuzzy c-means) to identify call types and assess their discreteness. We additionally used Uniform Manifold Approximation and Projection (UMAP) to visualize the separation of pulses using both extracted features and spectrogram representations. Supervised approaches showed low inter-observer reliability and poor classification accuracy, indicating that pulse types were not discrete. We propose an updated pulse classification approach that is highly reproducible across observers and exhibits strong classification accuracy using support vector machines. Although the low number of call types suggests long calls are fairly simple, the continuous gradation of sounds seems to greatly boost the complexity of this system. This work responds to calls for more quantitative research to define call types and quantify gradedness in animal vocal systems and highlights the need for a more comprehensive framework for studying vocal complexity vis-a-vis graded repertoires.

animal behavior and cognition↗