Not All Bees Pollinate: Using Machine Learning to Predict Ecological Roles

Image credit: José Neiva Mesquita-Neto

Understanding Pollination Through Ecological Function

Pollination is one of the most important ecological processes on Earth, but understanding how different insects contribute to it remains surprisingly difficult. While many bee species visit flowers, they do not all transfer pollen. Distinguishing between different ecological functions has traditionally required labor-intensive field experiments that are difficult to scale.

At our co-organized 2025 NeurIPS Workshop on AI for Non-Human Animal Communication, Alef Iury Siqueira Ferreira shared how machine learning can help shift the focus from identifying which species visited a flower to understanding the ecological role it played. You can see the paper here.

Drawing on expertise in ecology and machine learning, the team developed acoustic models that classify bees according to their ecological function, offering a scalable way to study pollination.

In this Community Spotlight, Alef discusses the challenges of measuring pollination in the field, why function-based monitoring may provide deeper ecological insights than species inventories alone, and how machine learning could help researchers better understand the complex relationships between pollinators, plants, and the ecosystems they support.

What challenges do researchers typically face when trying to measure pollination quality in the field?

While many insects visit flowers, they play different ecological roles. Some transfer pollen between flowers, contributing to plant reproduction, while others primarily collect nectar or pollen without transferring pollen in the same way. Simply observing that an insect visited a flower, or even identifying it taxonomically, does not directly tell us how much it contributed to pollination. In practice, measuring pollination quality often requires biologically grounded validation, such as single-visit pollen deposition experiments, which are labor-intensive, time-consuming, and difficult to scale under real field conditions. This is why assessing pollination in the field remains challenging: the real question is not only who visited the flower, but what ecological function that visitor actually performed.

How did collaboration between ecologists and machine learning researchers shape this project?

The collaboration between ecologists and machine learning (ML) researchers was the fundamental driver of this project, integrating expertise in computer science, pollination ecology, and bioacoustics to solve complex agricultural challenges. This partnership allowed for a multi-layered approach to automated bee monitoring that neither field could achieve in isolation.

The ecologists (Catholic University of Maule – UCM, Chile) provided consolidated experience in pollination ecology, specifically vibration-pollination in tomato and blueberry crops. They managed fieldwork, conducted traditional taxonomic identification, and measured the actual pollination function of various bee species. On the machine learning side (researchers from the Universidade Federal de Goiás – UFG, Brazil), contributed expertise in exploratory data analysis, acoustic preprocessing, feature representation, and the construction of a robust classification pipeline. This included the use of data augmentation strategies together with strong classifiers, including deep learning approaches such as CNN- and Transformer-based models, to improve robustness under real field conditions, even in the presence of class imbalance and the limited size of the dataset.

The ecologists motivated a significant shift in the project’s goal: moving from merely identifying bee species to recognizing their ecological function. Because taxonomic identity is often an indirect link to the actual delivery of pollination services, ecologists helped machine learning researchers focus on distinguishing “true pollinators” (which increase fruit size and yield) from “ineffective pollinators” (which do not aid pollination). This functional grouping provides a much clearer picture of ecosystem services than species-level inventories.

A bumblebee collecting nectar from white blueberry flowers on a branch.
Image credit: José Neiva Mesquita-Neto

What steps did you take to ensure the model could handle real-world field recordings with background noise?

Unlike traditional methods that use noise removal or attenuation, we choose to retain all environmental noise (such as traffic, birds, wind, and human speech) in the training data. This decision was based on the principle that training a model on noisy data allows it to generalize better when it encounters similar “unavoidable” noise during real-world testing. At the same time, we also considered the opposite risk: when background noise is preserved, the model may learn to rely on external cues that are not actually related to bee flight or sonication sounds. In that case, instead of learning the biologically relevant acoustic patterns, the model could overfit to recording-specific artifacts, which would harm generalization.  To reduce this risk, we used dynamic data augmentation during training to make the classifiers more robust to variability in real-world audio recordings.

“Unlike traditional methods that use noise removal or attenuation, we choose to retain all environmental noise (such as traffic, birds, wind, and human speech) in the training data. This decision was based on the principle that training a model on noisy data allows it to generalize better when it encounters similar “unavoidable” noise during real-world testing.”

How did you deal with class imbalance between common and rare pollinator types in the dataset?

In our work, we addressed the inherent class imbalance in the bee buzzing datasets – where common species naturally have more recorded samples than rare ones – through a combination of careful data splitting, data augmentation, and appropriate evaluation metrics.

The data splitting strategy was designed to preserve both biological and acoustic consistency. We used a stratified split based on species to ensure that each functional group and bee species was proportionally represented across the training, validation, and test sets. In addition, all segments extracted from the same original recording were kept within the same split, which helped avoid leakage of recording-specific information across splits and reduced the risk of biased evaluation. We also used data augmentation during training to increase the diversity of the audio samples and improve robustness, especially under limited and imbalanced data conditions. These techniques included masking parts of the acoustic representation, exposing the model to different portions of the same audio over training, and linearly combining pairs of audio samples to create more varied training examples. Together, these strategies helped the classifiers become more robust to variability in real recordings.

Finally, we found that classifying bees by functional roles (e.g., “effective” vs. “ineffective” pollinators) instead of individual species demonstrated greater robustness to overfitting. This is because functional groups concentrate more samples into fewer classes, which facilitates higher model “hit” probabilities compared to species-level inventories.

Figure 1: The bee buzzing recordings were segmented into flight and floral sonication sounds, augmented, and divided into training and testing sets. A pretrained convolutional neural network (CNN) was then fine-tuned to distinguish between the two pollination functional groups based on their acoustic signals. Finally, the model was evaluated using unseen recordings to assess its predictive performance.

What does this work suggest about moving from species-based monitoring to function-based monitoring in ecology?

Our work suggests that moving from species-based monitoring to function-based monitoring offers a more scalable, ecologically meaningful, and practically useful alternative for assessing ecosystem services. While traditional taxonomy is essential for biodiversity assessments, the sources indicate that a bee’s taxonomic identity is only an indirect link to the actual delivery of pollination services.

“Our work suggests that moving from species-based monitoring to function-based monitoring offers a more scalable, ecologically meaningful, and practically useful alternative for assessing ecosystem services.”

The functional classification focuses on the actual ecological roles organisms play within their communities. Traditional species identification is often time-consuming, expensive, and limited by a global shortage of expert taxonomists. A function-based approach allows non-experts, such as farmers and agronomists, to evaluate pollination quality without needing specialized taxonomic training. 

Also, classifying ecological functions can be technically more effective for AI. Models trained on functional groups demonstrate greater robustness to overfitting compared to species-level inventories. This is because functional categories concentrate more samples into fewer classes, which increases the data density per class and improves the model’s predictive probability.

Directly recognizing the contribution of specific bees to crop income can motivate farmers to adopt practices that support the most effective pollinators. This promotes a shift toward sustainable agriculture by highlighting the economic value of native bee communities.

How could automated recognition of pollinator function change how farmers or ecologists monitor crop pollination?

Automating the recognition of pollinator function represents a shift from identifying “who” is visiting a flower to understanding “what” that visitor is actually doing for the crop. This transition could fundamentally change monitoring by moving beyond traditional species-level inventories to direct assessments of ecosystem services.

Traditionally, ecologists have used flower visitation as a proxy for pollination, but this is often inaccurate because not all visitors are effective pollinators. Moreover, manual monitoring is labor-intensive, often requiring researchers to spend hundreds of hours in the field to observe and capture individuals.

Automated systems can now distinguish “true pollinators” from “ineffective pollinators”. Since effective pollination increases fruit set and enhances fruit size and weight, automated recognition allows for a direct prediction of crop quality and quantity.

The farmers and agronomists could use smartphone-compatible applications to identify the pollination roles of visiting bees in real-time. Understanding which local bees are effective pollinators, could help farmers create habitats and management practices that support those species.

A smiling man with curly hair and a beard, wearing a dark sweater, standing outdoors with greenery and parked cars in the background.

Alef Iury Siqueira Ferreira received the B.S. degree in Computer Science from the Federal University of Goiás (UFG), Goiânia, Brazil, in 2022. He is currently pursuing the M.Sc. degree in Computer Science at the Federal University of Goiás, where he has been involved in collaborative research projects with private companies and universities worldwide. His research interests include deep learning, representation learning, speech synthesis, and bioacoustics.

Disclaimer: Our Community Spotlight series highlights research and perspectives from across the animal communication community. Featuring this work does not imply endorsement by Earth Species Project, but reflects our commitment to sharing emerging research and advancements across the field.

Discover more from Earth Species Project

Subscribe now to keep reading and get access to the full archive.

Continue reading