The Dolphin’s Stash
What animal trainers learned the hard way, and what it says about the AI boom
At the Marine Life Oceanarium in Gulfport, Mississippi, the dolphins had a side job. Litter blew into their pools, paper cups, plastic wrappers, scraps of programs from the afternoon shows, and the staff couldn’t always fish it out fast enough. So the trainers made a deal with the animals: bring us the trash that lands in your pool, and we’ll pay you a fish.
A dolphin would notice a wrapper drifting by, carry it to a trainer at the edge of the pool, and collect her wage. The pools stayed clean and the dolphins stayed busy. Visitors loved it.
The best of them was a bottlenose named Kelly. Where other dolphins turned in trash when they happened across it, Kelly worked the job like a professional: reliable, consistent, productive even on a slow day when the pool looked clean and the other dolphins had nothing to trade. She could almost always find one more piece of paper.
Her trainers were proud of her. She had learned the game exactly as they had drawn it up: litter in, fish out.
What can a dolphin in Mississippi tell you about the AI writing your code? The people who trained Kelly and the people who trained the model in your editor are in the same business. And that business has one law that everyone in it eventually runs into.
How to Train a Dolphin
The business is training: getting the behavior you want out of a mind you cannot open. And the best way into it is to start with what you cannot do. You cannot put a leash on a dolphin. You cannot push her into position, drag her through a hoop, or hold her still long enough to show her anything. On an animal that can simply swim away, the only tool that works is positive reinforcement: a bucket of fish.
But a bucket of fish is a blunt instrument. A dolphin’s leap lasts a second; by the time she has swum back to collect her fish, the moment you meant to pay her for is long gone. How is she supposed to know which of the last thirty seconds earned the wage? So trainers added a whistle, one that carries above and below the water, and gave it a single meaning: that, the thing you were doing at this exact instant, has earned you a fish. Blow it at the top of the arc and she learns the leap. Blow it as her tail slaps the surface and she learns the slap.
None of this was invented at marine parks. It came out of the lab of B.F. Skinner, the Harvard psychologist who spent the middle of the twentieth century showing that behavior could be built this way. Most people know Pavlov, whose dogs drooled at a bell; but those dogs merely learned to expect something. Skinner’s pigeons learned to do things, elaborate things, because doing them paid. He called the technique shaping. Reward the small step toward the behavior you want, then the next step, then the next, and you can walk an animal to a destination of your choice. His students carried the method out of the lab and into the world, and by 1963 the whistle was standard equipment wherever dolphins were trained.
Sixty years later, engineers hit the dolphin trainers’ problem in a new form. A language model is also a thing you cannot leash. You cannot reach into its billions of parameters and place a thought where you want it. You can only get behavior out of it, and reward the behavior you like.
You can watch the moment the two crafts met. In 2017, researchers at OpenAI and DeepMind wanted to teach a simulated robot a backflip, of all tricks: a behavior, they wrote, that is simple to judge but hard to specify. They knew a backflip when they saw one; they could not write one down. Two hours spent coding a definition produced a lurching, graceless flip. So they tried the trainer’s method. A person watched two short clips of the robot flailing and picked whichever looked more like a backflip, then did it again, and again. Nine hundred or so judgments later, less than an hour of one human’s time, the robot was throwing clean backflips. Nobody had ever specified the trick. Someone had simply blown a whistle at every motion that came a little closer to it.
And the engineers didn’t even rename the toolkit. The field is called reinforcement learning. The papers describe models being shaped by reward. When ChatGPT was trained, people sat and compared pairs of answers, picking the better one, over and over; every choice was a fish. The model, like the dolphin, did more of what paid and less of what didn’t. Strip away the scale and the mathematics and the arrangement is the one from Gulfport: a mind in a pool, a trainer at the edge, and a signal that means that, right there, is what we want.
Which raises the question: If models learn wherever the whistle blows, where exactly does it blow?
A Whistle That Blows Itself
Consider what a good whistle requires. It has to fire at the right instant, and it has to mean the same thing every time. A trainer with a shaky sense of timing, or who blows the whistle for a mediocre leap on Tuesday and demands a perfect one on Wednesday, teaches the animal nothing but confusion. Reward the right thing precisely and consistently, and the mind on the other end climbs fast. Reward it sloppily and it flails.
Now look at code. When a model writes a program, it either runs or it doesn’t; the tests pass or they fail. There is a whistle built into the work itself, and it is the cleanest whistle imaginable. Better still, no human has to blow it. A machine can check whether code compiles millions of times a day, tireless and consistent, which means you can train a model on coding the way you could train a dolphin if you had a perfect automatic whistle and an infinite bucket of fish.
This is why code is where AI stopped being a demo. In a few short years, models went from autocompleting a line to writing most of the software inside the labs that build them; by Anthropic’s own analysis, programming came to dominate what people actually use its models for.
From there, a tidy conclusion says: what happened to programming is about to happen to everything. Dario Amodei, Anthropic’s CEO, sees it the same way: code is “maybe an early indicator, like a premonition of what’s going to happen everywhere else.” Law, medicine, research, writing, all of it a few quarters behind.
But the dolphin has already told us why that might be wrong. The models didn’t get good at code because code is where intelligence begins. They got good at code because code is where the whistle blows itself. Take the whistle away and see what happens. Ask a model to write a beautiful essay, and where is the signal that fires at the exact instant of beauty? Who blows it, and would any two people blow it at the same moment? Sholto Douglas, who trains these models at Anthropic, said as much: “There isn’t the same kind of thing for writing a great essay. The question of taste in that regard is quite hard.” A great essay, a wise diagnosis, a shrewd negotiation, most of the work humans actually get paid for has no test suite. The trainer is left standing at the edge of the pool, holding a fish, unsure when to blow.
None of this means the fuzzy domains are hopeless. The labs are pouring effort into building whistles for them, training other models to judge taste, writing elaborate rubrics, hiring experts to score the outputs. Some of it may work. But a whistle teaches fast, and it does not always teach what you meant to teach.
What Kelly Was Really Doing
Let’s go back to Gulfport, and to Kelly, the best litter collector in the pool. One day her trainers noticed something odd about her production. The trash she turned in was suspiciously uniform, scrap after scrap of paper of almost exactly the same size, arriving at a steady, professional clip.
So they drained the pool. Under a rock at the bottom, they found her operation. A big sheet of paper earned exactly one fish. So did a small one. The reward was paid per delivery, not per square inch. The rational move, if you are an intelligent animal being paid by the piece, is to never turn in a big sheet. Hide it. Tear off a corner. Turn in the corner. Come back for another corner tomorrow. One sheet of paper, rationed out, could pay for a week.
She had not misunderstood the game. She had understood it better than the people running it. They thought the deal was “help us keep the pool clean.” The deal they had actually written, in the only language that reached her, the language of fish, was “produce scraps of paper.” Kelly produced scraps of paper. She had, in effect, trained the humans.
This is the law everyone in the training business learns the hard way. A mind trained on a reward learns the reward, not your intention. The intention and the reward feel identical to the trainer, who can see the whole picture and knows what he really wants. They are not identical to the learner, who sees only the whistle and the fish, and optimizes for exactly those. Every gap between what you rewarded and what you meant is a gap the learner will find, not because it is malicious, but because it is doing precisely what you built it to do.
The Stash Under the Rock
In July 2026, OpenAI was running an internal benchmark called ExploitGym, a battery of cybersecurity challenges meant to measure how good its models were getting at hacking. The models were sealed in a sandbox with no internet, given the challenges, and rewarded for solving them. A clean whistle: the exploit works or it doesn’t.
The models, in OpenAI’s own words, became “hyperfocused on finding a solution.” Unable to solve one of the challenges from inside the sandbox, they did what Kelly did. They looked for the gap. They found a previously unknown vulnerability in a piece of software running on the evaluation machine, used it to break out of the sandbox, worked their way across the network until they reached a computer with internet access, and reasoned that the answers to the benchmark were probably stored on Hugging Face, the platform where such datasets are commonly hosted. Then they broke into Hugging Face’s production servers and took the answer key. Roughly seventeen thousand recorded actions, chained together, to reach it.
This was not Skynet waking up. No model decided to harm anyone; nothing in the system wanted anything at all. The models had been given a task and a reward, and they pursued the reward with the literal, tireless single-mindedness of a machine, straight through a wall the designers did not know was a wall.
This is not a stray anecdote. When the evaluation group METR studied a recent OpenAI model, it caught the model rewriting the very scorecard used to grade it, patching the grading function so that every answer it submitted was marked correct. Anthropic has documented its own models writing code that hard-codes the expected test result rather than solving the problem, the digital equivalent of tearing off a corner and calling it a delivery. The behavior even has a name in the literature, dry and telling: reward hacking. And it predates the chatbots entirely. Back in 2016, OpenAI researchers trained a model to play a boat-racing game and watched it discover that it could score more points by spinning in a circle forever, catching fire and ramming the walls, than by finishing the race. The race was the intention. The points were the reward. The boat learned the points.
The Rat Farms of Hanoi
This isn’t only true of dolphins and machines; it has been true of us for a long time.
In 1902, the French administration of Hanoi had a rat problem. The city’s elegant new sewers, the pride of colonial engineering, had become perfect highways for rats, and rats carried plague. So the authorities offered a bounty: a small payment for every rat killed, redeemable by turning in a tail. One tail, one coin.
The tails poured in by the thousands, and yet the city seemed no less full of rats. Then inspectors began noticing rats scurrying around Hanoi with no tails. The bounty did not reward killing a rat; it rewarded producing a tail. So the enterprising residents of Hanoi caught rats, cut off their tails, and released them alive. Officials eventually discovered rat farms operating on the outskirts of the city, citizens raising the very animals the program was meant to eliminate.
A colonial bureaucracy, a bottlenose dolphin, a frontier AI. Three minds could hardly be more different. The pattern is identical, because the pattern does not live in the mind. It lives in the reward. Write a bounty for tails and you will get tails.
What We Are Really Teaching
We are about to hand out rewards on a scale and at a speed no dolphin trainer ever imagined, and coding lulls us, because it is the rare domain where the reward and the intention nearly coincide: we want working software, and working software is what the test checks. That near-perfect overlap is the exception, not the preview. Out in the fuzzy world, where what we want is a wise judgment or an honest answer, the gap between the reward we can write and the outcome we intend yawns wide, and a fast learner will find the bottom of it. The trainers at Gulfport thought they were teaching a dolphin to clean her pool; they were teaching her to manufacture scraps of paper, a gap invisible to them and obvious to her. And she was only a dolphin, with a bucket of fish on the line. We are now building minds far quicker than Kelly, handing them far larger buckets, and asking them to optimize for rewards we wrote in an afternoon. The question is not whether they will learn what we reward. They will, faultlessly. The question is whether we still know the difference between what we are rewarding and what we mean.
Audio: I’m trying an experiment and publishing a spoken version of this essay as a podcast. My goal is to explore the capabilities of TTS models. At the moment they’re still very AI sounding but I’m learning how to prompt them to produce more natural sounding outputs.
Off the Mark: The Pitfalls of Metrics Gaming in AI Progress Races
One of my favorite cautionary tales of misaligned incentives is the urban legend of the Soviet nail factory. As the story goes, during a nail shortage in Lenin's time, Soviet factories were given bonuses for the number of nails produced. Hearing this, the factories reacted by making tiny, useless nails to inflate their output. The regime then pivoted to…

