bioRxiv ScienceSearch

Biology subjects

Jones, K. E.

Publications and source records attributed to Jones, K. E..

4 recordsLinked to original sources

CityNet - Deep Learning Tools for Urban Ecoacoustic Assessment

O_LICities support unique and valuable ecological communities, but understanding urban wildlife is limited due to the difficulties of assessing biodiversity. Ecoacoustic surveying is a useful way of assessing habitats, where biotic sound measured from audio recordings is used as a proxy for biodiversity. However, existing algorithms for measuring biotic sound have been shown to be biased by non-biotic sounds in recordings, typical of urban environments.\nC_LIO_LIWe develop CityNet, a deep learning system using convolutional neural networks (CNNs), to measure audible biotic (CityBioNet) and anthropogenic (CityAnthroNet) acoustic activity in cities. The CNNs were trained on a large dataset of annotated audio recordings collected across Greater London, UK. Using a held-out test dataset, we compare the precision and recall of CityBioNet and CityAnthroNet separately to the best available alternative algorithms: four acoustic indices (AIs): Acoustic Complexity Index, Acoustic Diversity Index, Bioacoustic Index, and Normalised Difference Soundscape Index, and a state-of-the-art bird call detection CNN (bulbul). We also compare the effect of non-biotic sounds on the predictions of CityBioNet and bulbul. Finally we apply CityNet to describe acoustic patterns of the urban soundscape in two sites along an urbanisation gradient.\nC_LIO_LICityBioNet was the best performing algorithm for measuring biotic activity in terms of precision and recall, followed by bulbul, while the AIs performed worst. CityAnthroNet outperformed the Normalised Difference Soundscape Index, but by a smaller margin than CityBioNet achieved against the competing algorithms. The CityBioNet predictions were impacted by mechanical sounds, whereas air traffic and wind sounds influenced the bulbul predictions. Across an urbanisation gradient, we show that CityNet produced realistic daily patterns of biotic and anthropogenic acoustic activity from real-world urban audio data.\nC_LIO_LIUsing CityNet, it is possible to automatically measure biotic and anthropogenic acoustic activity in cities from audio recordings. If embedded within an autonomous sensing system, CityNet could produce environmental data for cites at large-scales and facilitate investigation of the impacts of anthropogenic activities on wildlife. The algorithms, code and pre-trained models are made freely available in combination with two expert-annotated urban audio datasets to facilitate automated environmental surveillance in cities.\nC_LI

ecology

Bat Detective - Deep Learning Tools for Bat Acoustic Signal Detection

O_LIPassive acoustic sensing has emerged as a powerful tool for quantifying anthropogenic impacts on biodiversity, especially for echolocating bat species. To better assess bat population trends there is a critical need for accurate, reliable, and open source tools that allow the detection and classification of bat calls in large collections of audio recordings. The majority of existing tools are commercial or have focused on the species classification task, neglecting the important problem of first localizing echolocation calls in audio which is particularly problematic in noisy recordings.\nC_LIO_LIWe developed a convolutional neural network (CNN) based open-source pipeline for detecting ultrasonic, full-spectrum, search-phase calls produced by echolocating bats (BatDetect). Our deep learning algorithms (CNN FULL and CNN FAST) were trained on full-spectrum ultrasonic audio collected along road-transects across Romania and Bulgaria by citizen scientists as part of the iBats programme and labelled by users of www.batdetective.org. We compared the performance of our system to other algorithms and commercial systems on expert verified test datasets recorded from different sensors and countries. As an example application, we ran our detection pipeline on iBats monitoring data collected over five years from Jersey (UK), and compared results to a widely-used commercial system.\nC_LIO_LIHere, we show that both CNNFULL and CNNFAST deep learning algorithms have a higher detection performance (average precision, and recall) of search-phase echolocation calls with our test sets, when compared to other existing algorithms and commercial systems tested. Precision scores for commercial systems were reasonably good across all test datasets (>0.7), but this was at the expense of recall rates. In particular, our deep learning approaches were better at detecting calls in road-transect data, which contained more noisy recordings. Our comparison of CNNFULL and CNNFAST algorithms was favourable, although CNNFAST had a slightly poorer performance, displaying a trade-off between speed and accuracy. Our example monitoring application demonstrated that our open-source, fully automatic, BatDetect CNNFAST pipeline does as well or better compared to a commercial system with manual verification previously used to analyse monitoring data.\nC_LIO_LIWe show that it is possible to both accurately and automatically detect bat search-phase echolocation calls, particularly from noisy audio recordings. Our detection pipeline enables the automatic detection and monitoring of bat populations, and further facilitates their use as indicator species on a large scale, particularly when combined with automatic species identification. We release our system and datasets to encourage future progress and transparency.\nC_LI

ecology

Role of inter-related population-level host traits in determining pathogen richness and zoonotic risk

Zoonotic diseases are an increasingly important source of human infectious diseases, and host pathogen richness of reservoir host species is a critical driver of spill-over risk. Population-level traits of hosts such as population size, host density and geographic range size have all been shown to be important determinants of host pathogen richness. However, empirically identifying the independent influences of these traits has proven difficult as many of these traits directly depend on each other. Here we develop a mechanistic, metapopulation, susceptible-infected-recovered model to identify the independent influences of these population-level traits on the ability of a newly evolved pathogen to invade and persist in host populations in the presence of an endemic pathogen. We use bats as a case study as they are highly social and an important source of zoonotic disease. We show that larger populations and group sizes had a greater influence on the chances of pathogen invasion and persistence than increased host density or the number of groups. As anthropogenic change affects these traits to different extents, this increased understanding of how traits independently determine pathogen richness will aid in predicting future zoonotic spill-over risk.

epidemiology

Evaluating Bayesian spatial methods for modelling species distributions with clumped and restricted occurrence data

1. Statistical approaches for inferring the spatial distribution of taxa (Species Distribution Models, SDMs) commonly rely on available occurrence data, which is often non-randomly distributed and geographically restricted. Although available SDM methods address some of these problems, the errors could be more directly and accurately modelled using a spatially-explicit approach. Software to implement spatial autocorrelation terms into SDMs are now widely available, but whether such approaches for inferring SDMs are an improvement over existing methodologies is unknown.\n\n2. Here, within a simulated environment using 1000 generated species ranges, we compared the performance of two commonly used non-spatial SDM methods (Maximum Entropy Modelling, MAXENT and Boosted Regression Trees, BRT) to a spatially-explicit Bayesian SDM method (Integrated Laplace Approximation, INLA), when the underlying data exhibit varying combinations of clumping and geographic restriction. Finally, we tested whether any recommended methodological settings for all methods were further impacted by spatially non-random patterns in these data.\n\n3. Spatially-explicit INLA was the most consistently accurate method, being most or equal most accurate in 5 out of 8 data sampling scenarios. Within high-coverage sample datasets, all methods performed fairly similarly, but when sampling points were randomly spread BRT had a 1-3% greater accuracy over the other methods and when samples were clumped, spatial-INLA had a 4%-8% better in AUC score. Alternatively, when sampling points were restricted to a small section of the true range, all methods were on average 10-12% less accurate, with higher variation among the methods. None of the recommended settings for the different methods were found to be sensitive to clumping or restriction of data, except the complexity of the INLA spatial term.\n\n4. INLA-based modelling approaches can be successfully used to account for spatial autocorrelation in an SDM context and, by taking account of random effects, produce outputs that can better elucidate the role of covariates in predicting species occurrence. Given that it is often unclear what the drivers are behind data clumping in an empirical occurrence dataset, or indeed how geographically restricted these data are, spatially-explicit INLA-based SDMs may be the better choice when modelling the spatial distribution of target species.

ecology