bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.07.03.187252

Deep learning and computer vision will transform entomology

Abstract

ABSTRACTMost animal species on Earth are insects, and recent reports suggest that their abundance is in drastic decline. Although these reports come from a wide range of insect taxa and regions, the evidence to assess the extent of the phenomenon is still sparse. Insect populations are challenging to study and most monitoring methods are labour intensive and inefficient. Advances in computer vision and deep learning provide potential new solutions to this global challenge. Cameras and other sensors that can effectively, continuously, and non-invasively perform entomological observations throughout diurnal and seasonal cycles. The physical appearance of specimens can also be captured by automated imaging in the lab. When trained on these data, deep learning models can provide estimates of insect abundance, biomass, and diversity. Further, deep learning models can quantify variation in phenotypic traits, behaviour, and interactions. Here, we connect recent developments in deep learning and computer vision to the urgent demand for more cost-efficient monitoring of insects and other invertebrates. We present examples of sensor-based monitoring of insects. We show how deep learning tools can be applied to the big data outputs to derive ecological information and discuss the challenges that lie ahead for the implementation of such solutions in entomology. We identify four focal areas, which will facilitate this transformation: 1) Validation of image-based taxonomic identification, 2) generation of sufficient training data, 3) development of public, curated reference databases, and 4) solutions to integrate deep learning and molecular tools.Significance statement Insect populations are challenging to study, but computer vision and deep learning provide opportunities for continuous and non-invasive monitoring of biodiversity around the clock and over entire seasons. These tools can also facilitate the processing of samples in a laboratory setting. Automated imaging in particular can provide an effective way of identifying and counting specimens to measure abundance. We present examples of sensors and devices of relevance to entomology and show how deep learning tools can convert the big data streams into ecological information. We discuss the challenges that lie ahead and identify four focal areas to make deep learning and computer vision game changers for entomology.Competing Interest StatementThe authors have declared no competing interest.GlossaryBin pickingan industrial term for robots that pick up one of many objects randomly placed in a container.Convolutional Neural Network (CNN)a deep learning algorithm in the family of neural networks with serval different layers commonly applied for image recognition and classification. A CNN can be trained to recognize various objects and patterns in an image. There are four main different operations in a CNN: convolution, activation functions, sub sampling, and fully connected layer. During training the learnable parameters of each convolutional and fully connected layer are adjusted so the CNN is able to recognize different patterns of the training data and used for final image classification.Data augmentationa technique that can be used to artificially expand the size of a training dataset by creating modified images with objects of interest for classification.Machine learninga subset of artificial intelligence associated with creating algorithms that can change themselves without human intervention to get the desired result – by feeding themselves through structured data.Deep learninga subset of machine learning where algorithms are created and function similarly to machine learning, but where there are many levels of these algorithms, each providing a different interpretation of the data it conveys.DNA barcodingIdentification of a species using a short, standardised gene fragment.Initializationdescription of an object to be tracked.Training dataclassified images (e.g. images of known species identified by experts) that are recorded to train a deep learning model.Precisionthe number of true positives divided by the sum of true positives and false positivesRecallalso called the true positive rate, is the number of true positives divided by the sum of true positives and false negatives.Classification accuracythe sum of true positives and true negatives divided by the total number of specimens.View Full Text

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hoye, T. T., Arje, J., Bjerge, K., Hansen, O. L., Iosifidis, A., Leese, F., Mann, H., Meissner, K., Melvad, C., Raitoharju, J.. 2020-07-04. Deep learning and computer vision will transform entomology. https://doi.org/10.1101/2020.07.03.187252

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Beyond Single-Metric Assessments: Uncovering Masked Butterfly Declines via Multi-Scalar Analysis in Central Alberta

1. This study analyzed 21 years (2000-2025) of butterfly count data from Central Alberta, integrated with intensive 5-year (2021-2025) high-resolution intra-seasonal sampling. 2. Long-term macro-scale analysis revealed a significant decline in Shannon Diversity, a change that remained obscured when relying solely on traditional metrics of species richness and evenness. 3. This diversity decline was primarily driven by the severe, long-term collapse of the native Common Ringlet (Coenonympha tullia). 4. Four other dominant species--Cabbage White (Pieris rapae), Clouded Sulphur (Colias eriphyle), European Skipper (Thymelicus lineola), and Common Wood Nymph (Cercyonis pegala)--maintained long-term population stability, though their abundances were significantly constrained by extreme winter minimum temperatures and rapid spring warming. 5. High-resolution intra-seasonal analysis (2021-2025) demonstrated that community indices and species-specific abundances were strongly limited by daily weather, particularly wind velocity and temperature. 6. These findings illustrate that while traditional metrics like richness and evenness are fundamental to community ecology, they provide incomplete insights when applied in isolation; they are most effective when utilized as part of a complementary, multi-scalar framework. 7. This study highlights the necessity of coupling multi-decadal historical datasets with high-frequency, fine-scale sampling to accurately identify the mechanisms of community turnover that simpler metrics may overlook. 8. The results underscore the critical importance of standardized citizen science monitoring in quantifying environmental impacts and establishing conservation priorities for terrestrial insect groups.

ecology

From concentration to export: resource contrasts and bee traits shape pollinator spillover to crops

Floral plantings can either concentrate bees or export them to adjacent crops, yet the ecological conditions influencing these outcomes remain unclear. Here, we develop a mathematical model as proof of concept for our previous integrative hypothesis: concentrator and exporter outcomes can arise as alternative, context-dependent outcomes of the same underlying resource-selection process. Using bees as a model and focusing specifically on spillover from floral plantings to crops, we identified resource-specific thresholds separating concentration- and export-favoring conditions. Our model translates differences in relative patch attractiveness into context-dependent concentration and export outcomes and generates resource-specific, testable predictions about the conditions favoring pollinator movement into crops. In our simulations, the concentrator-exporter transition occurred at a lower flowering-intensity contrast than at pollen or nectar contrasts, which suggests that flowering intensity may provide an initial cue for bee movement, whereas nectar and pollen rewards refine or sustain bee responses once crops are perceived as attractive. Spillover thresholds differed among resource contrasts, whereas response steepness varied across bee-trait and community scenarios. Under the model's trait-sensitivity formulation, predicted spillover probability responded more strongly to flowering contrast for specialists than for generalists; colony size amplified this response, whereas bee richness dampened it. Together, these patterns show how flowering and resource contrasts interact with bee traits and community context to shape predicted spillover. Our results confirm that the concentrator and exporter hypotheses can be understood as context-dependent outcomes of the same ecological process rather than as mutually exclusive alternatives. Experimental tests of the predicted thresholds conducted in the field could reveal when and where floral plantings are most likely to promote bee spillover to crops, potentially supporting crop pollination.

ecology

A Computational Re-evaluation of Spatial Trials for Zoonotic Tuberculosis Control: Model Misspecification, Diagnostic Miss-classification, and the Illusion of Wildlife Culling Efficacy

1. Wildlife reservoir management frequently relies on the Randomised Badger Culling Trial's (RBCT) trade-off hypothesis, which posits that reductions in cattle herd infections are offset by a perturbation effect driven by disrupted host dispersal. This paper evaluates the computational and epidemiological robustness of this historical trial, which serves as the foundational empirical experiment guiding zoonotic tuberculosis (Mycobacterium bovis) control policies. 2. Using generalized linear mixed models with a generalized Poisson error distribution to explicitly address historical data overdispersion, this study contrasts traditional parametric inference against exact cluster-constrained permutation tests across distinct operational definitions of disease incidence. 3. Non-parametric diagnostics reveal that previously reported treatment and perturbation effects render as statistical artifacts under exact non-parametric permutation. Inside culling zones, parametric significance fails to withstand exact permutation verification due to extreme data leverage in localized cluster blocks. 4. Crucially, when diagnostic misclassification biases are eliminated by analysing total reactor datasets, all apparent culling effects disappear, and information criteria overwhelmingly favour nested null architectures. Unconfirmed reactors likely represent true biological infections missed by low-sensitivity post-mortem macro-necropsy, proving that host removal tracks observation noise rather than genuine zoonotic transmission pathways. 5. Finally, empirical scaling conducted in this study identifies a novel mathematical saturation effect, demonstrating that this sub-linear scaling is an operational artifact of unmodelled herd-level disease recurrence over time. 6. Policy implications. Because current zoonotic tuberculosis intervention frameworks are built upon a structurally misspecified statistical model, they have driven large-scale veterinary policies resulting in substantial, unevidenced ecological and economic interventions while failing to provide genuine public health, animal health, or disease control benefits.

ecology