Raymond UzwyshynIdeas · Research · Artificial Intelligence
Science, Research & Discovery

The Dolphins and AI: How AI is learning to listen to the sea's most sophisticated speakers

On a humid April morning in 2025, Laela Sayigh sat in her Woods Hole Oceanographic Institution office, headphones clamped over her ears, listening to sounds that might rewrite our understanding of animal minds. Not…

Cover graphic for The Dolphins and AI: How AI is learning to listen to the sea's most sophisticated speakers

On a humid April morning in 2025, Laela Sayigh sat in her Woods Hole Oceanographic Institution office, headphones clamped over her ears, listening to sounds that might rewrite our understanding of animal minds. Not the signature whistles she'd spent four decades studying—those individualized acoustic autographs that dolphins use like names—but something stranger, more systematic, more suggestive. Something that whispered of words.

Fifty-four years ago, when researchers first began systematically recording the bottlenose dolphins of Sarasota Bay, Florida, the notion that machines might one day decode animal communication belonged firmly to science fiction. Today, as transformer architectures and self-supervised learning algorithms reshape what's computationally possible, Sayigh's discovery arrives at the convergence of two exponential curves: one measuring marine mammal behavioral data accumulated across generations, the other measuring silicon's capacity to find patterns in that data. The question is no longer whether AI can analyze dolphin whistles—deep learning classifiers already achieve 95% accuracy. The question is whether what it finds constitutes something we can legitimately call language.

The Names They Carry

The dolphins of Sarasota Bay have been watched, photographed, and acoustically monitored since 1970, creating what is surely the most intimate portrait of any wild animal society. Approximately 170 individuals inhabit these waters—six generations tracked with near-obsessive thoroughness, their relationships mapped, their movements charted, their voices archived. It is the Stanford longitudinal study of wild cetaceans, and Sayigh has devoted her professional life to listening.

Her doctoral work in the mid-1980s confirmed what Melba and David Caldwell had first suspected: each dolphin invents a unique frequency-modulated whistle during its first year of life—a sonic signature as distinctive as a fingerprint, yet fundamentally different. Fingerprints are accidents of embryology. Signature whistles are choices.

"Think of it," Sayigh says, leaning back in her chair, "a year-old dolphin creating its own name. Not inheriting it. Not having it assigned. Choosing it through vocal production learning, shaping sound through trial and error until it settles on a pattern it will carry for decades." She pauses. "That's not just unusual among animals. That's essentially unprecedented."

The properties that make signature whistles extraordinary can be enumerated: they're individually distinctive, referentially stable across decades, socially transmitted, and—most remarkably—used as labels. Dolphins copy each other's signature whistles at low rates, effectively calling specific individuals by name. Playback experiments confirm they can identify particular dolphins from signature whistles even when researchers strip away all "voice" characteristics, leaving only the frequency modulation pattern—pure information, divorced from individual timbre.

But names alone do not make a language. Proper nouns label; they don't describe, don't explain, don't argue. For decades, researchers wondered whether dolphin vocal communication stopped there—a society of individuals announcing themselves, calling to each other across the acoustic vastness of Sarasota Bay, but never saying more than "I am here."

Then Sayigh started listening, really listening, to everything else.

The Background Hum Becomes Foreground

Non-signature whistles—NSWs in the taxonomic shorthand—had always been there, comprising roughly half of all dolphin whistle production. Researchers had noticed them, noted them, generally ignored them. They were the linguistic equivalent of background noise, the acoustic furniture of dolphin sociality. Individual-specific signature whistles were the story; NSWs were mere margin notes.

Sayigh's insight, born from sheer persistence and pattern-recognition honed across thousands of hours of recordings, was to approach NSWs with systematic rigor. What if these weren't individualized at all? What if they were shared?

The Sarasota Dolphin Whistle Database—926 recording sessions from 293 individually identified dolphins, some tracked across 43 years—provided the temporal depth and individual specificity necessary for the analysis. Sayigh's team developed algorithms to extract whistle contours, measure similarity, cluster types. The computational work was painstaking: separating signal from noise in underwater recordings where waves, boats, and shrimp compete for acoustic bandwidth; matching whistles to identified individuals; accounting for variability in production.

The results were worth the wait. Twenty-two distinct non-signature whistle types emerged, produced by multiple individuals across the Sarasota community. Two types stood out with particular clarity:

NSWA, produced by at least 25 different dolphins, consistently elicited avoidance behavior in playback experiments. Not flight exactly, but withdrawal, caution, a collective drawing-back. The researchers hypothesized: alarm signal, danger notification, the cetacean equivalent of shouting "heads up."

NSWB, produced by at least 35 individuals, appeared in contexts of novelty and uncertainty—when dolphins encountered unfamiliar objects, unexpected situations, the cognitive experience of "what is this?" The team ventured: query signal, perceptual question mark, the acoustic manifestation of curiosity.

Here was the pattern that mattered: arbitrary sound patterns, associated with consistent behavioral contexts, used across a population as shared communicative units. Here, possibly, were words.

The Statistical Substrate of Meaning

At NeurIPS 2025—the Neural Information Processing Systems conference that has become machine learning's annual pilgrimage—a workshop on December 6th brought together researchers who are teaching computers to parse the vocalizations of creatures whose throats evolved along radically different trajectories than ours. The convergence was explicit: Earth Species Project researchers, MIT computational linguists, engineers from Google and Meta, field biologists who smell of seawater and sunscreen, all gathering in San Diego's convention center to discuss how silicon might serve as Rosetta Stone for biological communication.

Oisin Mac Aodha, a computer scientist at the University of Edinburgh with the kind of Scottish accent that makes "algorithm" sound like poetry, demonstrated BatDetect2—deep learning infrastructure for identifying bat echolocation calls. His system achieves 0.88 mean average precision across seventeen UK bat species, processing ultrasonic recordings in real-time on Raspberry Pi devices small enough to mount on trees. "The technical problem," he explained, "isn't just classification accuracy. It's building annotation tools where human expertise and machine learning bootstrap each other, where each improves the other iteratively."

Julie Elie, a Berkeley neuroscientist whose work bridges birdsong and primate vocalization, presented evidence that zebra finches possess semantic categories for their calls—not just acoustic categories (calls that sound similar) but meaning categories (calls that signify similar things). When birds make errors in call classification tasks, they confuse different contact calls with each other more often than they confuse contact calls with alarm calls, despite acoustic differences that would predict the opposite pattern. "That's categorization by meaning," Elie emphasized. "That requires mental representation of what calls mean, not just what they sound like."

The methodological consensus was striking: self-supervised learning on unlabeled audio, leveraging the distributional structure of vocalizations themselves. You don't need to know what whistles mean to find patterns in how they're used. You don't need translation to discover grammar.

A team from MIT presented WhaleLM—effectively a "sperm whale language model"—trained to predict whale vocalizations from conversational history. Sperm whale codas, those rhythmic click patterns that echo through ocean depths, exhibit order dependence and long-range dependencies spanning up to eight codas. The model could predict current whale behavior with 72% accuracy and future actions with 86% accuracy from coda sequences alone. Not perfect prediction, but far better than chance—evidence that vocalizations encode information relevant to behavioral coordination.

If patterns in whale clicks correlate with behavior, and patterns in finch calls encode semantic categories, and patterns in dolphin whistles cluster into types associated with consistent contexts... what are we seeing? What are machines helping us hear?

Saussure in the Salt Water

Ferdinand de Saussure, the Swiss linguist who died in 1913 without publishing the book that would bear his name, never contemplated dolphins. His focus was human language—specifically, how to think about it structurally, systematically, scientifically. His insight, assembled from student lecture notes after his death, was that language consists of signs: arbitrary linkages between sound patterns (signifiers) and mental concepts (signified).

The arbitrariness is crucial. The English word "tree" bears no natural relationship to trees themselves—it's a convention, shared among English speakers, linking a particular sequence of phonemes to the concept TREE. French uses "arbre," German "Baum," Japanese "ki"—different signifiers, same signified, proof that the connection is conventional rather than natural.

Dolphin signature whistles are Saussurean signs par excellence. The frequency modulation pattern identifying individual F146—a dolphin Sayigh has recorded for decades—bears no acoustic resemblance to F146's physical form, behavioral tendencies, or social position. It's an arbitrary signal, meaningful only because the Sarasota community treats it as meaningful. That's the essence of symbolic reference: not mimicry or icon, but convention.

The newly discovered NSWs extend this pattern. The acoustic structure of NSWA doesn't sound like danger; dolphins don't whistle in frequencies that somehow iconically represent shark attacks or boat traffic. Yet the whistle consistently correlates with avoidance behavior across dozens of individuals. Signifier (NSWA's frequency pattern) linked to signified (danger, threat, withdrawal) through shared convention.

But Saussure offered another crucial distinction: langue versus parole, the abstract linguistic system versus individual speech acts. Langue is the shared code, the collective competence; parole is actual utterance, individual performance. Human children acquire langue through exposure to parole, extracting systematic patterns from the variable flux of actual speech.

Could Sarasota Bay dolphins possess something analogous to langue? A shared repertoire of signature whistles and NSW types that individual dolphins instantiate through particular vocalizations? The consistency of NSWA across 25 producers and NSWB across 35—stereotyped patterns maintained across individuals—suggests exactly this. Not individual expression, but participation in a collective code.

Umberto Eco, the Italian semiotician who moved effortlessly between medieval manuscripts and James Bond novels, defined semiotics more broadly: the study of everything that can be used to lie. By "lie" he meant the capacity to represent what is not immediately present—the fundamental property of symbols. When a dolphin produces another individual's signature whistle to call them, that dolphin is representing someone absent, engaging in displaced reference, one of Charles Hockett's famous "design features" of language.

If NSWA functions as an alarm signal, it represents danger that may not be immediately visible—abstract reference to a category of threatening situations. That's symbolic thought, not stimulus-response. That's the difference between perceiving threat and communicating about threat.

The Architecture of Attention

The technical capability undergirding these discoveries rests on three architectural innovations that have transformed what's computationally tractable.

Convolutional neural networks—those layered systems of artificial neurons that learn hierarchical features from data—proved their worth first. A 2024 study by the Sarasota research team applied MobileNetV2, a CNN architecture optimized for mobile devices, to signature whistle classification. The network achieved 95.8% accuracy identifying individual dolphins from their whistles under clean conditions, dropping only to 92.6% under realistic background noise. Data augmentation—artificially creating training examples by stretching time, shifting pitch, mixing with ocean ambient sound—proved critical for robustness.

But CNNs process fixed-duration inputs, struggling with sequences of variable length and long-range dependencies. Enter transformers, the architecture that conquered natural language processing and is now being adapted for bioacoustics. The innovation is attention: rather than processing sequences left-to-right linearly, attention mechanisms let models attend to relevant parts regardless of position. For dolphin vocalizations, this means capturing how whistles relate across conversational time, how preceding calls constrain subsequent ones, how meaning might emerge from sequential structure.

The frontier, though, is self-supervised learning. Traditional machine learning required labeled data: human experts annotating thousands of recordings, marking "this is a signature whistle," "this is NSWA," "this is F146." Self-supervised models learn from unlabeled audio by predicting masked portions—given this whistle's beginning, what frequency comes next? Given these three calls, what's the fourth? The model learns acoustic representations without explicit labeling, extracting structure from distribution.

AVES—Animal Vocalization Encoder based on Self-supervision, developed by Earth Species Project—learns acoustic representations by predicting masked audio regions, then transfers those representations to diverse bioacoustic tasks. The remarkable finding: models pretrained on human speech transfer effectively to animal sounds. Apparently, general acoustic structure—the ways sounds unfold in time, the statistical regularities of vocalization—transcends species boundaries.

For dolphin research, this means: train a foundation model on hundreds of hours of unlabeled Sarasota Bay recordings, let it discover latent structure, then fine-tune on the much smaller set of expert-annotated examples. The model learns from distribution what individual biologists could never hear—patterns across thousands of vocalizations, temporal correlations spanning decades, structural regularities invisible to human attention.

The Body Beneath the Signal

But here's the rub, the place where enthusiasm meets epistemic humility: pattern is not meaning. Statistical structure is not semantics. A language model trained on English text can generate grammatical sentences without understanding anything—John Searle's famous Chinese Room argument demonstrated this conceptually; GPT demonstrates it practically.

The same risk haunts bioacoustic AI. A transformer trained on dolphin whistles might learn to predict what whistle typically follows NSWA, might generate synthetic whistles conforming to distributional patterns, might even successfully classify novel recordings. None of this proves the system understands what the whistles mean to dolphins.

Meaning, the embodied cognition theorists argue, is grounded in sensorimotor experience. Humans understand "warmth" through the sensation of temperature on skin, "up" through gravitational orientation and muscle proprioception, "red" through the subjective experience of seeing red things. Linguistic concepts aren't free-floating abstractions; they're anchored to bodies moving through environments.

Dolphins inhabit bodies radically different from ours, moving through environments we experience only as alien visitors. Their vocal apparatus lacks cords entirely—sound production occurs in nasal passages via phonic lips that vibrate pneumatically. The melon, that bulbous fatty organ dominating the forehead, acts as acoustic lens, focusing sound into directional beams. They hear partly through their lower jaw, which conducts sound to inner ear structures. They can produce clicks and whistles simultaneously from different sides of their nasal system—a capacity with no human equivalent, no human phenomenology.

Consider echolocation: dolphins perceive their environment primarily by analyzing reflected sound, building mental representations from acoustic structure. They "see" by listening, perceiving size, shape, texture, internal composition from echo patterns. A world interpreted through time-of-flight calculations, frequency-dependent absorption, complex acoustic scattering—phenomena humans access only through technological mediation, never through direct experience.

Marine biologist Diana Reiss describes it vividly: "Imagine if you could see through walls by humming, if your voice could paint pictures, if sound gave you the internal structure of objects. That's dolphin perception. Their communication system is embedded in that perceptual world."

Three-dimensional sociality compounds the difference. Unlike terrestrial mammals constrained to surface movement, dolphins inhabit unbounded volumetric space where group members routinely disperse across square kilometers then rejoin. This fission-fusion social structure—group composition changing constantly across hours—creates extreme demands for individual recognition and relationship maintenance. Shark Bay dolphins maintain social networks of 60-70 regular associates, comparable to the largest primate societies, but spread across a space where visual monitoring is impossible.

Signature whistles and shared NSWs function within this embodied reality. When NSWA triggers avoidance behavior, the meaning of that signal is grounded in echolocation-based spatial perception, aquatic predator-avoidance kinematics, and social relationships tracked through acoustic contact alone. The whistle doesn't mean "danger" in some abstract propositional sense; it means something embodied—move this way, not that way; attend here; coordinate departure.

Jakob von Uexküll, the Estonian biologist who coined the term "Umwelt" in the 1930s, emphasized that each species inhabits a species-specific perceptual world. The dolphin Umwelt includes ultrasonic frequencies up to 150 kHz (humans max out around 20 kHz), echolocation-derived spatial representations, social information encoded in acoustic patterns we perceive only through technological prosthesis. AI can identify patterns in dolphin vocalizations without inhabiting dolphin Umwelt. It can learn the code without accessing the meaning.

Or, more precisely: it can learn correlations without causation, associations without understanding, prediction without comprehension.

The Question of Questions

Which brings us to the central question, the one that haunts every attempt to decode animal communication: What would it mean to truly understand what dolphins are saying?

Not to classify their whistles—we can do that. Not to predict their sequences—transformers manage that respectably. Not even to correlate vocalizations with behavior—Sayigh's playback experiments demonstrate consistent relationships. But to know, really know, what information those whistles carry for their producers and receivers. To access dolphin meaning, not just human interpretations of dolphin patterns.

Charles Hockett's famous design features—his attempt to specify what distinguishes human language from animal communication—included properties like displacement (talking about things not present), productivity (generating novel utterances), duality of patterning (meaningless sounds combining into meaningful units). Dolphin signature whistle copying demonstrates displacement. The 22 NSW types hint at combinatorial structure. But do dolphins exhibit true productivity, generating novel "sentences" from existing "words"? Do they recursively embed, the hallmark of syntactic language?

We don't know. The honest answer, the scientifically responsible answer, is: we don't know.

What we do know is this: dolphins evolved vocal learning independently from humans—convergent evolution of the rare capacity to modify vocalizations based on auditory feedback. They possess individually distinctive learned signals that function referentially and are used as labels. They produce shared, stereotyped non-signature whistles associated with consistent behavioral contexts. Their vocalizations exhibit sequential structure with long-range dependencies. They maintain complex social networks requiring sophisticated individual recognition and relationship tracking.

That's not nothing. That's actually quite a lot. Whether it constitutes "language" in the technical linguistic sense matters less than understanding what it actually is—this sophisticated, flexible, learned communication system evolved by large-brained, long-lived social mammals inhabiting three-dimensional acoustic space.

Sayigh puts it plainly: "I don't need dolphins to have human language to find their communication fascinating. I need to understand dolphin communication on its own terms. That's what makes it scientifically interesting—not confirmation that they're like us, but discovery of what they actually are."

Listening Forward

The morning I visited Sayigh's lab, she pulled up a spectrogram on her monitor—frequency on the vertical axis, time on the horizontal, sound visualized as shapes. "This," she said, pointing to a rising-falling contour, "is FB11's signature whistle. I recorded him for the first time in 1984. He's still alive, still producing this exact pattern. Forty-one years of carrying the same name."

She clicked to another file. "This is his mother's whistle. You can see where he incorporated elements—he learned from her but made it his own. That's vocal production learning. That's culture in the technical sense: information transmitted socially across generations, modified through learning."

Another click. "And this—this is NSWA. Six different individuals produced this over the last decade, all in contexts where we documented boat approaches or predator presence. Same whistle, different dolphins, consistent context."

She leaned back. "Now imagine we get funding to array the whole bay with synchronized hydrophones. We track every whistle from every dolphin over five years, correlate with behavior, social context, environmental conditions. Feed that to self-supervised learning systems that extract distributional patterns we'd never notice. What might we find?"

It's a good question, one that 54 years of patient observation and the sudden acceleration of AI capability has brought within reach. The data exists—nearly a million whistles archived, attributed to known individuals, contextualized through parallel behavioral observation. The models exist—transformer architectures, self-supervised encoders, few-shot learners that work with limited labeled data. The experimental frameworks exist—playback protocols for testing dolphin responses to synthetic stimuli.

What's missing is the synthesis: bringing all these pieces together in systematic pursuit of understanding. Not just classifying whistles but mapping their distributional structure. Not just finding patterns but testing whether dolphins respond to those patterns in ways consistent with proposed meanings. Not just building models but using models to generate hypotheses testable through experiment.

The Earth Species Project, that nonprofit organization funding much of this convergence between bioacoustics and AI, has an ambitious vision: decode animal communication across species, create tools for two-way interspecies communication, fundamentally reshape human relationships with the natural world. It's the kind of grand ambition that attracts both enthusiasm and skepticism, supporters and critics, funding and scrutiny.

But whether or not that maximal vision succeeds, something genuine is happening at this intersection. The computational tools have caught up to questions that field biologists have been asking for decades. The temporal depth of datasets like Sarasota Bay has reached the threshold where machine learning can extract meaningful structure. The theoretical frameworks—from Saussure to Sebeok, from embodied cognition to biosemiotics—provide scaffolding for interpreting what models reveal.

There's a moment in science when multiple exponential curves converge, when data accumulation, computational capacity, theoretical understanding, and funding interest all spike simultaneously. Protein folding had such a moment when AlphaFold launched. Exoplanet detection had it when Kepler's data met machine learning.

Animal communication is having it now.

The Sound of Thinking

On my last evening in Woods Hole, Sayigh invited me out on the research vessel for a twilight run across Sarasota Bay. The water was glass-smooth, reflecting pink-orange clouds. We deployed hydrophones off the stern, and within minutes, the boat's speakers crackled with clicks, whistles, the acoustic texture of dolphin sociality.

"You hear that?" Sayigh asked, pointing to a whistle rising and falling across three seconds. "That's FB11 again. Same individual we saw on the spectrogram this morning. He's forty-one years old now, probably calling to his current coalition partners."

Another whistle answered, a different frequency pattern. "That's BB03, his primary associate. They've been together for over two decades. They know each other's signatures as well as humans know each other's faces."

The clicks intensified—echolocation, dolphins scanning the boat, us, the hydrophones we'd deployed. "They're investigating," Sayigh said. "Trying to figure out what we are, what we're doing. That's curiosity, right there. That's active problem-solving about novel stimuli."

A longer, more complex whistle emerged from the speakers. Sayigh tensed slightly, listening with the focused intensity of someone who's devoted a lifetime to this particular frequency range. "NSWB," she said quietly. "Query signal. Right on cue—they encountered something unexpected, and they're communicating about it."

We sat in silence, listening. The dolphins continued vocalizing, weaving patterns of sound through darkening water. Clicks for perception, whistles for communication, the constant acoustic monitoring of social and physical space.

"You asked earlier," Sayigh said after a while, "whether I think dolphins have language. Here's what I think: they have something. Something sophisticated, something learned, something that carries information between complex minds. Whether 'language' is the right word depends on how you define language. But they definitely have something worth understanding on its own terms."

She gestured toward the speakers. "And for the first time—really for the first time—we have tools that might help us understand it. Not translate it, maybe. Not speak it ourselves. But understand its structure, its organization, its relationship to their lives. That's what keeps me listening after all these years. The possibility that they might finally help us hear."

The boat rocked gently. The dolphins continued their conversation, patterns in the darkness, signals we're only beginning to parse. Somewhere in Silicon Valley, transformer models trained on half a million whistles were learning to predict sequences. Somewhere in Edinburgh, BatDetect2 was processing ultrasonic recordings in real-time. Somewhere at Berkeley, Julie Elie was mapping the neural pathways that enable vocal learning in mammals.

And in Sarasota Bay, under a sky turning from orange to purple, dolphins were talking. Not to us—we're still mostly irrelevant to their communicative lives, occasional curiosities at the margins. They were talking to each other, as they have for millions of years before humans waded into the water and started listening.

The difference now is that we're finally learning to hear. Not perfectly, not completely, but better than we ever have before. The dolphins carry on their conversations, patterns in saltwater and time, while we—blessed with silicon and algorithms and five decades of patient observation—inch closer to comprehension.

Whether we'll ever truly understand remains an open question. But the listening has begun in earnest, and that may be enough.

This is a two part article based on a NeurIPS AI/Interspecies Communication Workshop. This is the narrative contextualization regarding the workshop. There is also a more scientifically oriented version describing dolphins and workshop discoveries available here: https://www.linkedin.com/pulse/deciphering-dolphin-communication-ai-deep-learning-uzwyshyn-ph-d--axocc

#AIforAnimalCommunication #NeurIPS2025 #DolphinResearch #Bioacoustics #MarineBiology #MachineLearning #NeuralSymbolicAI

Originally published January 1, 2026. View the original publication ↗