What if a Model Knows More Than it Can Say?

Image credit: Gregoire Dubois / Flickr

Using Model Merging to Access What Bioacoustic Models Already Know

Bioacoustic foundation models are opening up new ways to study animal communication across species, including species they have never encountered during training. But what determines how much of what these models have learned we can actually access?

When Emanuele Rossi and Donato Crisostomi began experimenting with NatureLM-audio, they found that its performance could vary depending on how a question was phrased. This suggested that NatureLM-audio had learned useful information about animal sounds that it was not always able to express reliably through different linguistic instructions. Emanuele and Donato wanted to find out whether they could improve that flexibility while preserving the model’s bioacoustic capabilities.

In research presented at the AI for Non-Human Animal Communication workshop at NeurIPS 2025, they explored model merging as a way to do this. They combined NatureLM-audio with the language model it was originally built from, improving its ability to follow varied instructions while retaining its expertise in animal sounds. 

When tested on species from taxonomic families it had never encountered during training, the merged model achieved a more than 200% relative improvement in F1 score, setting a new state of the art for this challenging evaluation setting. 

We’re still using the learnings from this paper in the latest version of NatureLM-audio 2.0, to be released soon! 

In this Community Spotlight, Emanuele shares what he and Donato learned through model merging, why generalization across species matters, and how studying what these models learn could help uncover patterns in animal communication across species.

What motivated this research, and what limitation in current bioacoustic foundation models were you trying to address?

Broadly, my research focuses on building machine learning models that generalize across species and investigating what they can teach us about animal communication. As part of that, we started playing with NatureLM-audio soon after its release, as it was an exciting new tool for bioacoustics and we wanted to explore what it could do.

We quickly noticed something surprising: its performance depended heavily on how a question was phrased. When we asked the model to identify the species in a recording, it performed very well when prompted for the common name of the species and slightly less well when asked for the scientific name. But when we asked for both names together, its performance dropped dramatically (see Figure below), even though the underlying bioacoustic task was essentially unchanged.

Infographic showing prompts for audio species identification: Common Name, Scientific Name, and Combined Prompt, with a bar graph comparing accuracy results for Watkins and CBI.
Figure 1: Same task, different phrasing. We asked NatureLM-audio to identify the species in recordings from two datasets: Watkins (marine mammals) and CBI (birds). Each time we used one of the three prompts shown on the left. The model is accurate when asked for the common name or the scientific name alone. When asked for both names in a single answer, accuracy collapses to 10% and 3%, even though identifying the species is just as hard.

This suggested that the limitation was not simply a lack of knowledge about animal sounds, but rather that this knowledge could not always be accessed and expressed reliably through different linguistic instructions.

Further experiments showed that this was part of a broader brittleness in the model’s language understanding and instruction-following, an important limitation for a tool intended to support bioacousticians and ethologists. Researchers should be able to ask questions naturally, combine instructions, and explore their data without having to discover the exact prompt formulation that produces a useful answer. This motivated us to investigate whether we could restore the model’s linguistic flexibility without sacrificing its bioacoustic expertise.

“Researchers should be able to ask questions naturally, combine instructions, and explore their data without having to discover the exact prompt formulation that produces a useful answer.”

Your research uses “model merging.” What does that mean in simple terms, and why did you think it could help NatureLM-audio?

Put simply, model merging combines two machine learning models to bring together capabilities that each has separately, without having to do any additional (and expensive) training. To see why this could be useful, imagine one model that is an expert in bioacoustics but cannot converse fluently in English, or any other language. It struggles to understand what you want, and you struggle to understand its answers. Another model speaks English fluently but knows little about bioacoustics. The goal of merging them is to create a model that can converse fluently with you about bioacoustics.

That was exactly our situation with NatureLM-audio! Our collaborator Donato Crisostomi, an expert in model merging, immediately saw an opportunity to merge it with a model that was much stronger at language. This is one of the things I love about working across disciplines: Donato does not normally work on animal communication, but his expertise turned out to be exactly what this problem needed.

And the solution turned out to be beautifully simple. The language model we merged it with was actually the one NatureLM-audio had been built on. Because of that relationship, we could blend the two by changing a single number in the code. At one end, we had the original model’s language abilities; at the other, NatureLM-audio’s bioacoustic specialization. We found a point in between where the model became much better at following varied instructions while retaining its expertise in animal sounds.

What changed after merging NatureLM-audio with its original language model?

The first effect was that the merged model became much better at following instructions while retaining its bioacoustic expertise. For example, when asked to provide both the common and scientific names of a species, accuracy increased from 6% to 45% on one dataset and from 12% to 63% on another. This confirmed that model merging could recover much of the linguistic flexibility lost during specialization.

“When asked to provide both the common and scientific names of a species, accuracy increased from 6% to 45% on one dataset and from 12% to 63% on another. This confirmed that model merging could recover much of the linguistic flexibility lost during specialization.”

What excited me most was that the improvement extended beyond the prompts we had initially investigated. We evaluated the merged model on recordings of species from taxonomic families that had not appeared during training. The model received a list of 40 possible species and had to identify the correct one directly from the recording, without any additional training or examples. On this task, its F1 score increased from 0.09 to 0.28, representing more than a 200% relative improvement and establishing a new state of the art for this challenging evaluation setting.

Our analysis suggested that NatureLM-audio’s audio encoder already contained useful representations of these unfamiliar species. The original model could often distinguish between the sounds, but it struggled to follow the instruction to select one of the permitted labels and instead produced answers outside the list. Merging improved its adherence to the prompt, allowing it to make better use of knowledge that was already present within the model.

A prompt for closed-set classification, asking for the common name of a species mentioned in an audio clip.
Bar chart showing error rates for 'Out-of-set' predictions (orange) and 'In-set but wrong' predictions (yellow) across different scaling factors ranging from 0.0 to 1.0.
Figure 2: Where the model’s mistakes come from. We asked the model to identify the species in a recording by choosing from a fixed list (prompt shown above), while blending NatureLM-audio with its original language model. The scaling factor controls the blend: 0 is closest to the original language model, and 1 is the unmodified NatureLM-audio. Each bar shows the error rate, split into two kinds of mistakes. Orange marks answers outside the permitted list, a failure to follow the instruction. Yellow marks answers from the list that name the wrong species, a failure of bioacoustic knowledge. Unmodified NatureLM-audio (right) rarely picks the wrong species from the list, but it often ignores the list entirely. Toward the language-model end (left), the model sticks to the list but picks wrong more often. In between, around 0.4, both kinds of error are low.

This was particularly exciting because the intervention was designed primarily to restore linguistic flexibility, yet it also produced a major improvement in generalization to unseen species. It showed that better instruction-following is not only a matter of making a model easier to interact with, but it can also determine whether the model is able to express and apply its bioacoustic capabilities successfully.

Why is it important to develop models that are generalizable to species they haven’t encountered during training?

One of the most exciting things about NatureLM-audio was that it could perform tasks involving species it had never encountered during training. That matters enormously in bioacoustics. There are far too many species for us to collect and label enough recordings to train a separate model for each one.

Imagine placing a recorder in a forest. Ideally, you would want a tool that can help you analyze the sounds it captures even when some of the species were absent from its training data. If the model only works for species it has seen before, much of that recording may remain out of reach.

Scale is one of the greatest things machine learning can bring to the study of animal communication: build a tool once, then use it across many species and recordings, rather than developing a new system or manually analyzing the data each time. But generalization also raises a scientific question: what has a model learned about animal sounds that allows it to work across species?

What does this work tell us about the capabilities already contained within bioacoustic foundation models like NatureLM-audio and how we might make better use of them?

Our results suggest that NatureLM-audio could distinguish sounds from unseen species better than its initial performance led us to believe. Some of that ability was already there; the way we prompted the model, and how well it could follow those instructions, affected whether we could make use of it.

But the open questions interest me even more. As these models are trained on more species and more recordings, what patterns are they learning about animal sounds? Do those patterns reflect differences the animals themselves perceive? And do any of them hold across species?

That is the connection I want to explore: developing better models, then investigating what they learn to generate hypotheses about animal communication.

What do researchers have to gain if foundation models are able to perform flexibly for a variety of prompts and instructions?

Flexible instruction-following could make a foundation model feel less like a collection of narrow classifiers and more like a research assistant. Imagine something like ChatGPT or Claude, but with deep expertise in bioacoustics (and cheaper). Researchers could ask questions naturally, reformulate them, combine instructions, and request answers in the format that best suits their analysis, without having to find the precise prompt that works for each task.

It could also make the bioacoustic knowledge learned by these models useful across a wider range of tasks. The same model might help detect vocalizations, classify species or individuals, compare recordings, and describe acoustic patterns without requiring a separately trained system for each question. That broad utility is part of the vision behind NatureLM-audio, but it depends on the model being able to follow varied instructions and sustain a conversation about the data. Finally, this would make such tools more accessible to bioacousticians and ethologists who do not have extensive machine learning expertise.

“Researchers rarely know every question they want to ask in advance, and interesting findings frequently lead to new questions. A model that can respond reliably to varied instructions could support this exploratory process.”

More broadly, scientific research is often iterative. Researchers rarely know every question they want to ask in advance, and interesting findings frequently lead to new questions. A model that can respond reliably to varied instructions could support this exploratory process, helping researchers move from large collections of recordings to patterns and hypotheses that can then be examined and validated through biological expertise and further experiments.

What could generalizable models reveal about patterns in animal communication across species, bringing us closer to understanding how they communicate?

One concrete example is vocal identity. One line of my research asks whether machine learning can tell individual animals apart from their vocalizations alone, including individuals the model has never heard before, and whether that ability transfers across species. The key question for me is not just whether a model can tell individuals apart, but how it does so: which acoustic features does it use, and do those features also matter to the animals themselves?

This connects to a broader question about what can be shared across species. As a model is trained on more species and tasks, learning patterns that apply across them is often more efficient than learning each case separately. Large language models on human languages offer a good example: a multilingual model need not learn every concept independently in every language: a fish lives underwater whether we describe it in English, Italian, or Spanish. Research suggests that such models develop some shared representations alongside features specific to each language, such as vocabulary. This shared structure can then be leveraged to, for example, translate between languages even with few examples. 

Obviously the analogy has limits: the communication systems of different species are vastly more diverse than human languages, and different animals perceive and inhabit the world (their umwelt) in very different ways. We cannot assume that a pattern learned across species represents a shared meaning, but if there are shared communication primitives, i.e. recurring features in how animals produce, combine, or respond to sounds, a model trained across species might help us find them.

“A model distinguishing two animals does not by itself show that it uses the same information as the animals. It gives us something concrete to investigate together with biologists. That loop between AI and scientific discovery is what I find most exciting.”

Finding a pattern in a model is just the beginning of the scientific work. We then need to investigate what the model has learned, ask whether it reflects something biologically meaningful, and test the resulting hypotheses against animal behavior. In the case of vocal identity, a model distinguishing two animals does not by itself show that it uses the same information as the animals. It gives us something concrete to investigate together with biologists. That loop between AI and scientific discovery is what I find most exciting.

An underwater scene featuring a sea lion swimming near a diver, surrounded by rocks and bubbles in clear blue water.
Image: Emanuele with a sea lion while diving in the Galapagos.

Emanuele Rossi is a postdoctoral researcher in the Gladia group at Sapienza University of Rome, focusing on developing novel machine learning methods to deepen our understanding of non-human communication. He’s particularly fascinated by what this can reveal about other forms of intelligence, and how such insights might reshape our relationship with the natural world.

His move into animal communication grew from a lifelong love of animals and a desire to use machine learning to understand how other species experience and interact with the world. It also continues a broader interest in finding structure in complex systems, from graphs during his PhD at Imperial College and Twitter, to structural biology at Vant AI.

Earlier in his career, he was part of Fabula AI, a deep learning startup for fake news detection that was acquired by Twitter. He holds degrees in Computer Science from Imperial College London and the University of Cambridge. Outside research, he loves scuba diving, hiking, and quiet moments in nature. You can find more about his work at emanuelerossi.co.uk.

Disclaimer: Our Community Spotlight series highlights research and perspectives from across the animal communication community. Featuring this work does not imply endorsement by Earth Species Project, but reflects our commitment to sharing emerging research and advancements across the field.

Discover more from Earth Species Project

Subscribe now to keep reading and get access to the full archive.

Continue reading