BirdCODE: Scaling Zero-Shot Sound Event Detection for Bioacoustics

Image credit: Bejan Adrian / Pexels

Key Takeaways

  • BirdCODE is a supervised deep learning model that offers zero-shot sound event detection, enabling researchers to detect and classify bird vocalizations with temporal precision and without dataset-specific fine-tuning.
  • The model was developed using a combination of training strategies using weakly labelled recordings, synthetic soundscapes, and pseudo-labelled data to perform sound event detection and classification across more than 9,000 bird species.
  • Applying BirdCODE to more than one million recordings demonstrates how large-scale temporal annotation can enable new analyses of vocal behaviour across species, seasons, geography, and mixed-species interactions.
  • BirdCODE, the model weights, and large-scale detections are openly available on GitHub to support future research in Animal Language Processing.
  • This paper is a preprint and currently under review.

Sound event detection remains one of the central challenges in bioacoustics. While recent machine learning models have dramatically improved species classification, detecting the precise temporal boundaries of individual vocalizations remains difficult, particularly at large taxonomic scales where strongly labelled training data are scarce.

For researchers interested in animal communication, this limitation is significant. Studying communication at scale depends on accurate temporal annotation. Understanding interactions between individuals often requires knowing precisely when each vocalization begins and ends relative to every other vocalization. Solving this challenge is therefore a prerequisite for studying animal communication at scale.

Introducing BirdCODE 

To address this challenge, we developed BirdCODE (Bird Communication Detector), a deep learning model for large-scale sound event detection and classification across more than 9,000 bird species.

Unlike conventional sound event detection systems, BirdCODE does not rely on extensive strongly labelled training datasets. Instead, it learns from the vast collection of weakly labelled recordings contributed through citizen-science initiatives, combining them with synthetic soundscapes and pseudo-labelled predictions in a multi-stage training pipeline.

Figure 1: An overview of BirdCODE model training. Strongly labeled data are converted to per-frame species labels (7.6 Hz), which serve as the prediction target. On weakly labeled data (i.e. without onset and offset annotations), the model’s per-frame predictions are pooled into a single per-clip prediction for each species.

BirdCODE achieves state-of-the-art performance on both sound event detection and species classification benchmarks, extending sound event detection to roughly 100 times more bird species than previous bioacoustics-focused systems.

Figure 2: Performance comparison for BirdCODE against baseline methods using sliding-window classifiers. The number in parentheses below each model name is the number of WABAD datasets where the model achieved the top score in that metric. Histogram displays performance per recording site within WABAD (68 sites). Red line displays performance on held-out strongly labeled recordings from xeno-canto.

BirdCODE demonstrates that large-scale, zero-shot sound event detection is possible, even when strongly labelled datasets are unavailable. While strong benchmark performance is valuable, it is not the end goal. The real measure of a sound event detection system is whether it accelerates existing science and enables new discoveries. 

From Detection to Discovery

To demonstrate that potential, we applied BirdCODE to more than one million community-contributed recordings and used the resulting detections to investigate several biological questions. Anyone can load these detections on publicly licensed data from xeno-canto and iNaturalist within alp-data by following the instructions on the GitHub repo.

What We Found

Phylogenetic Patterns in Vocal Behaviour

Birds that are closely related often share similar traits in their vocalizations, a pattern shaped by their shared evolutionary history. Studying this connection has traditionally required extensive human annotation of individual vocalizations. Using BirdCODE detections instead, we examined whether vocalization duration tracked the avian phylogeny, drawing on recordings across a wide range of species and controlling for factors like body size, beak length, and habitat that might otherwise explain similarities between species.

Using BirdCODE detections, we found similarities in vocalization duration among closely related species across the avian phylogeny. This result complements prior research into how a bird’s evolutionary history influences its vocal behavior, and shows how BirdCODE can support new research in comparative bioacoustics.

Figure 3: Phylogenetic tree of birds with > 5 recordings containing detections. Exterior colors separate different phylogenetic orders, leaf color denotes average duration of vocalizations detected in that species, and interior edge color denotes ancestral state reconstruction.
Seasonal Variation in Vocal Behaviour
Image 1: European robin (Erithacus rubecula). Credit: Francis C. Franklin / CC-BY-SA-3.0
Image 2: Great Tit (Parus major). Credit: Andreas Trepte

Applying BirdCODE across year-round recordings of two common European songbirds, the great tit (Parus major) and the European robin (Erithacus rubecula) revealed seasonal shifts in vocal behaviour within species, including changes in call duration and acoustic characteristics in species such as the great tit and European robin. These matched known phenological patterns in these common songbirds, demonstrating BirdCODE’s ability to detect and quantify season variations at scale, without the intensive manual annotation that this kind of analysis typically requires.

Figure 4: Per-month variation in duration, dominant frequency, and spectral entropy of vocalizations by great tit (Parus major) and European robin (Erithacus rubecula). Shaded region is 95%
confidence interval.
Geographic Variation in Vocalizations
Image 3: Rufous-collared sparrow (Zonotrichia capensis costaricensis).
Credit: Charles J. Sharp / CC BY-SA 4.0

Analyses of BirdCODE detections also captured geographic differences in the songs of rufous-collared sparrows (Zonotrichia capensis), a small songbird widespread across South America. Their songs sometimes end in a trill or buzz, a feature used by some populations but not others, possibly linked to differences in habitat.

Figure 5: Geographic variation in spectral entropy of rufous-collared sparrow song. Left: color denotes average of recordings within 500km radius, and opacity decays with distance from recording location (transparent after 500km). Red circles denote 1000km radius around each city. Right: Spectrograms of songs sampled from near Quito (top shaded region, trill present) and Sao Paulo (bottom, trill present).

Mapping the average spectral entropy of these songs across a 500-km radius, we identified two well-sampled regions with contrasting patterns: around Quito, Ecuador, and around São Paulo, Brazil. Songs near Quito showed higher average spectral entropy, while those near São Paulo showed lower entropy.

Spectrogram 1: Rufous-collared sparrow song recorded near Quito, Ecuador, with trill. 64% of songs in this region have trills.
Credit: Charlie Vogt / CC BY-NC-SA 4.0
Spectrogram 2: Rufous-collared sparrow song recorded near near Sao Paulo, Brazil, without trill. No songs in this region have trills.
Credit: Fernando Igor de Godoy / CC BY-NC-SA 4.0

Since trills tend to sound noisier than non-trilled songs, we predicted that the higher-entropy region would contain more trills. A manual check of the recordings confirmed this: sparrows recorded near Quito often included a trill in their song, whereas none of the sparrows recorded near São Paulo did.

This demonstrates how large-scale sound event detection can support the study of geographic variation in vocal behaviour, including identifying region-specific dialects.

Figure 6: Top: Spectral entropy of songs with and without trills in each region. Bottom: 64% of songs from Quito had trills, whereas none from Sao Paulo had trills.
Cross-Species Vocal Interactions

In North American winters, chickadees, white breasted nuthatches, downy woodpeckers, and tufted titmice congregate into foraging flocks. Applying BirdCODE to recordings of these flocks, we looked at whether a vocalization from one species predicted a vocalization from another within the following five seconds, then compared that to what would be expected by chance. 

Figure 7. Weighted graph of cross-species vocal interactions; edge color and thickness correspond to deviation from null of how strongly vocalizations by one species predict those of the other. Stars indicate significant interactions (Bonferroni-corrected).

This revealed that chickadees respond to nuthatches and titmice more readily than expected by chance, and that white-breasted nuthatches respond similarly to downy woodpeckers. 

These results offer a scalable framework for investigating cross species interactions and coordination in natural soundscapes.

Building Infrastructure for the Future of Bioacoustics

At ESP, we build AI tools that help scientists overcome technical barriers and accelerate discovery. BirdCODE is one example of that approach. By releasing the model, code, and large-scale detections, we hope to provide infrastructure that researchers can build on in ways we can’t yet predict.

Discover more from Earth Species Project

Subscribe now to keep reading and get access to the full archive.

Continue reading