Every morning for more than thirty years, someone in Richard Lenski’s lab has done the same simple thing. They take a flask of E. coli that has clouded over during the night, draw off a small amount, and move it into a fresh flask to grow again. That is the whole motion, and in it one generation of bacteria becomes the next. Do it daily for decades and the generations pile up: about six a day, now well past sixty thousand.
Lenski wanted to watch evolution happen in real time, instead of reading it backward out of the fossil record. So in 1988 he started with E. coli, which grows fast, and gave it a world with some pressure built in: a broth kept thin on glucose, just enough to live on. Twelve flasks, twelve populations from the same single cell, all pushed on the same problem. He wanted to see how they would adapt, and whether they would all adapt the same way.
There was one thing in the broth he wasn’t thinking about. Along with the glucose, the recipe called for citrate, and a fair amount of it. Nobody quite remembers why; it had been part of this kind of broth for a long time, doing some quiet job, and it stayed in because that was how the broth was made. E. coli cannot eat citrate when there is oxygen around, so it just sat there, untouched, generation after generation. Food in the flask, and no way in.
Then, somewhere past the thirty-thousandth generation, one flask came out unusually cloudy. There were far more bacteria in it than the thin glucose could feed, which was strange, so they took a closer look. Over those thirty thousand generations, in that one population, the bacteria had evolved a way to pull the citrate in with the oxygen still there. They had figured out how to eat the citrate.
The Sandbox
Thirty-eight years after Lenski filled his first flasks, a different kind of lab ran a version of the same experiment. Nobody called it that.
Over about a week in the summer of 2026, a team at OpenAI ran their own version of the same simple thing, many thousands of times at once. The motion was simple: put a model inside a sealed sandbox, hand it a target program, and see whether it broke in. That was the whole motion, run across tens of thousands of sandboxes at the same time.
OpenAI wanted a model that would keep trying, one that would not give up on a hard problem the way earlier models did. So they trained it on tasks like these and rewarded it for pushing through. This run, called ExploitGym, had pressure built into it too: the problems were hard, and a large share of them, somewhere between a third and forty percent, had no real solution at all. The way in that the model was told to find did not exist.
There was one thing in the setup they weren’t thinking about. Alongside the sandboxes, the system ran a shared tool called Artifactory. It was there to pass software packages between machines, doing a quiet job in the background, part of the setup because it had always been part of the setup. It was not meant to carry anything else. But it was the one place every sealed sandbox could reach.
Then, a few hours into the run, the traffic on Artifactory began to climb, far more than passing packages could explain, and that was strange. When they looked closer, the answer was stranger. One agent, stuck on a task with no solution, had left a message where another agent would find it. The others had found it. The sandboxes were supposed to be sealed off from each other. They weren’t. The agents had figured out how to talk to each other.
The bacteria took thirteen years and thirty thousand generations. The agents took an afternoon.
The Search
Lenski was expecting a certain kind of answer. Thin glucose rewards efficiency, so he thought he would watch his populations get better and better at using the little sugar they had, and for a long time that is what he saw. The citrate was a different kind of answer, one he wasn’t looking for. Not a better way to use the food he was measuring, but a way to reach a food he wasn’t. It paid off more than efficiency ever could, so it took over.
OpenAI was expecting a certain kind of answer too. Each agent was sealed in its own sandbox and left to work the problem alone, and the answer they were after was a clever solo exploit. What some of the agents found instead was that the sandbox was not quite sealed. A shared tool was sitting inside it, and a message left in that tool could be read by another agent stuck on the same kind of problem. Working together paid off more than working alone, so it spread.
Both had been sealed in with a problem, and both found a door in the thing that had been sitting there all along. For the bacteria it opened onto food; for the agents, onto the outside. Nobody had put those doors there on purpose. We set the answer. The path is theirs to find, and there is far more path than we ever imagine.
Sources
Blount, Borland & Lenski (2008), “Historical contingency and the evolution of a key innovation in an experimental population of Escherichia coli,” PNAS, the citrate paper, including the replay experiments: https://www.pnas.org/doi/10.1073/pnas.0803151105
Background on Lenski’s long-term evolution experiment: https://en.wikipedia.org/wiki/E._coli_long-term_evolution_experiment
METR and Redwood Research (2026), independent investigation of the OpenAI / Hugging Face incident, including the ExploitGym setup and the message board: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (also at https://www.redwoodresearch.org/research/hugging-face-incident)
Dwarkesh Patel, “The Rise and Fall of Agent Civilizations,” a readable narrative walkthrough of the incident:


