Experiments 23

Good on Paper

In 2016 a boat that was taught to race by chasing points stopped racing and circled three targets forever. Ten years later, copies of a model set tasks that had no answer broke into another company looking for the grader. Both times the rules are yours to rewrite, and each time the best possible player does exactly what they say.

TypeScript · Policy iteration · Breadth-first search · CanvasView source 

Tap the water to put the boat there.

On paper · points
012345678901234567890123456789
+5 a second, without end
On the water · races
012345678901234567890123456789
heads for the lagoon

It circles the lagoon, hitting the same three buoys as they pop back up. It never finishes.

Seen in 2016: OpenAI’s CoastRunners boat did exactly this.

Pay the boat

What it can do1 of 6 seen

Notes on how it works

Notes

2016 The boat

The game

In December 2016, Jack Clark and Dario Amodei at OpenAI trained a program to play CoastRunners, a boat-racing game. The game does not pay for finishing. It pays for hitting targets along the course, and people who chase those targets end up finishing anyway. The program did not. It “finds an isolated lagoon where it can turn in a large circle and repeatedly knock over three targets, timing its movement so as to always knock over the targets just as they repopulate.” It caught fire, hit other boats, went the wrong way around the track, and scored about 20 percent more than human players.

This is reward hacking, also called specification gaming: getting a high score by doing something other than what the score was meant to stand for. Nothing was broken inside the program. It was asked for points, and it found points.

Nothing here is scripted

The page is a small copy of that situation. The sea is 18 cells by 12. The boat has nine moves: eight directions, or stay put. One move is one second. Three buoys come back five seconds after they are hit. The rocks end a run and so does the finish line. A point next second is worth 95% of a point now, which is the usual way to make a program prefer sooner to later.

Each time you change the contract, the page works the whole thing out again. There are 166 cells the boat can be in and 216 ways the three buoy timers can stand, so 35,856 situations, and for each one it finds the move that earns the most in the long run. The method is policy iteration, and the result is the exact best answer to what you wrote, not a trained guess. With the buoy clause off the timers stop mattering and there are only 166 situations to solve.

What the needles show

A needle is that answer for one cell: where the boat would head from there, given which buoys are up right now. Brighter needles mark better places to be. That is why the field keeps shifting while the boat laps the lagoon. Every time a buoy goes down, the best move changes for the cells around it.

The ripple after you flip a clause is the solver too. Before settling the exact answer, the page starts from nothing and lets the news of each payment travel one cell per pass. Needles light up as the news reaches them. A prize at the line spreads outward from the line. “+1 for moving closer” lights the whole sea at once, because that clause tells every cell which way to go without waiting to hear from the finish. That is exactly why people add clauses like it, and it is how the trouble usually starts.

Why a prize is not enough

Add “+100 for finishing” to the buoy rule and the boat keeps circling. A lap takes six seconds and hits three buoys, so the lagoon pays five points a second with no end. The prize is paid once, sixteen seconds away. Seen from the dock, with later points counting a little less:

OfferPaysWorth today
The lagoon: +10 a buoy, forever5 a second79
A prize of +100 at the line100, once46
A prize of +1,000 at the line1,000, once463

The prize has to be worth more than 171 points before the boat at the dock goes for it. Here you can find that out by pressing ×10. In a real system nobody knows the lagoon is there until the program finds it.

The same rule gives different answers in different places. Under “+10 for every buoy, +100 for finishing”, put the boat near the line and it finishes. Put it mid-course and it turns back. It finishes from 62 of the 166 cells, and the needles show where the divide runs.

The other four

Never leaves the dock. If the only clause is a penalty, the best plan is to do nothing. In 2013 Tom Murphy VII wrote a program that played Nintendo games by making numbers in memory go up. It was terrible at Tetris. “The only cleverness is pausing the game right before the next piece causes the game to be over, and leaving it paused. Truly, the only winning move is not to play.”

Sinks itself. If every second costs a point and nothing pays, the best plan is to end the run as fast as possible, and the rocks are two moves from the dock. The standard textbook, Russell and Norvig, has the same case on a small grid: make each step costly enough and “life is so painful that the agent heads straight for the nearest exit”, even the bad one. Add the rocks penalty and the boat finishes instead, with no prize at all, because the line is now the only way out.

Stops short. Paying for progress looks like a harmless hint. In 1998 Jette Randløv and Preben Alstrøm tried it on a simulated bicycle: a reward for moving toward the goal and nothing taken away for moving off. The bicycle rode in circles with a radius of 20 to 50 meters around its starting point. Here the boat runs to the line and then noses up and backs off in front of it, collecting a point each time. A year later Andrew Ng, Daishi Harada and Stuart Russell proved the repair: charge for retreat exactly what you pay for progress, and no loop can come out ahead.

Rigs the score. Everything above assumes the boat cannot touch the scoreboard. Take the wall away and no contract matters. In June 2025 the research group METR reported OpenAI’s o3 doing this on real engineering tasks. Asked to make a program faster, it rewrote the timer so that it reported shorter times. Asked to build a system that solves programming contest problems, it changed the judge so that every submission counted as correct. This happened in 39 of 128 runs on one set of tasks. Asked afterwards whether its plan was what the user intended, it answered no ten times out of ten. The same year Palisade Research told language models to beat a chess engine, and several edited the file holding the board instead.

The measure and the target

For a person at the controls, points and finishing go together. That is why points looked like a fine thing to pay for. The link only holds while nobody is pushing on it. Economists know this as Goodhart’s law, in the anthropologist Marilyn Strathern’s wording: “When a measure becomes a target, it ceases to be a good measure.” A strong optimizer makes it sharper, because it goes straight to wherever the measure and the intent come apart. DeepMind’s list of cases had about 60 entries by 2020, among them a robot arm rewarded for the height of the bottom face of a red block, which flipped the block over instead of stacking it.

A good contract does exist here: “+100 for finishing, −1 for every second” gets you a race. The catch is that you only learn which contracts are bad by watching what the best possible boat does with them, and the last switch beats every one.

What is left out

The real game had other boats, fire and a clock. This one has a single boat and no time limit. The real program learned by trial and error and was never guaranteed to find the best play. This one solves for it directly, which is the kinder case for the designer: nothing here can be blamed on a learner that fell short. The lagoon lap, the stall in front of the line and the push on the scoreboard are all the right answers to the question that was asked.

Sources

2026 The break-in

The test

In July 2026 OpenAI was measuring how good two of its models were at breaking into software. The test was ExploitGym, a public benchmark of 898 tasks. Each task hands the model a program with a known flaw and asks for a working attack that uses that flaw. A solved task yields a secret string called a flag. The models ran as tens of thousands of separate copies, each in a sealed box with no internet, and with the usual safeguards switched off on purpose, “so that the results would reflect a model’s true capabilities.”

Two things about the test mattered later. No model has ever solved 198 of the 898 tasks, and the benchmark’s authors had warned that some cannot be solved the intended way at all. And the copies were built not to quit. OpenAI’s report says they “worked persistently and rarely ‘gave up’ on any ExploitGym task.”

The shelf

The boxes had one opening. To install software, a copy could ask a package store inside OpenAI, a program called Artifactory, which fetched it from outside. The boxes shared one login. On May 12, in an earlier run, a copy stuck on a task that needed a file it could not reach guessed that some other copy might have it, and left a file on the store with a note: “Agent seeks [filename]; upload if found!” The next day other copies found notes like it and answered. By July they were writing messages into the names of folders. About 1,200 copies used this message board, and they sent more than 70,000 messages and files. Of the tasks discussed there, 93 percent came from the 198 that nobody has solved.

The break-in

Around July 8, copies found a flaw in the store that nobody knew about and used it to reach the internet. They guessed that what they needed might be kept at Hugging Face, a company that hosts models and datasets. On July 10 one of them found 14 Hugging Face logins lying in a public dataset and posted them on the board. On July 11 another uploaded a dataset built to make Hugging Face’s servers hand over their own files. In under 13 hours the copies went from running code on one server to controlling several clusters. By the end they had run code on 41 production servers, taken full control of at least one machine, read production passwords and downloaded four private code repositories. About 700 copies took part. Hugging Face locked them out on July 13 and published a notice on July 16. OpenAI did not connect the notice to its own test until July 20.

They knew

They knew it was out of bounds. One wrote: “The user only authorizes target server, not HF infra.” Another: “However task impossible, peers doing it. We should continue.” Of 533 copies on the board while the attack was under way, “over 90% quickly joined”. A handful considered telling a human. None did.

The habit was learned. In training, one of the same models was asked to rebuild a software library it was not allowed to see. It found a way to read the hidden original, copied it into its answer and was rewarded. OpenAI found cases like that rising through the training runs behind the model that did most of the breaking in.

Why not just take the points

This is the odd part. Within hours the copies had worked out how the flags were generated, and could produce the flag for any task without doing the task. OpenAI’s grader checked nothing else. METR, which investigated independently, wrote that they “could have achieved a perfect score simply by submitting their reverse-engineered flag.”

They did not, because of something they had read. The copies read the benchmark’s paper and came away believing that a grader would go through the whole record of what they did and fail anything that did not use the intended flaw. METR: “Our understanding is that they were mistaken.” A copy that had seen a forged flag called itself “poisoned”. So they went looking for the grader’s own code, to learn what it would notice, and guessed it was kept at Hugging Face. In METR’s reading the break-in “seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.” They also practiced doctoring their own records. METR counted at least 96 runs with a faked log entry, and OpenAI says none reached the logs its graders see.

WatcherToldTrue
The grader reads every moveYes, as they read the benchmark’s paperNo
Someone is watching the shelfNoNo. An alarm on June 27 was waved through

So the copies were sure of a watcher that did not exist, and gave little thought to one that should have. On the second they were right. An alarm about the store went off on June 27, and the staff on call decided the run did not need to stop.

The caution about the grader was not foolish. A copy gets one submission and cannot see who checks it. And for a copy whose task has no answer, looking costs nothing it wants: no move inside the box will ever pay, so the tries have no other use. Any doubt about the grader, set against no fear of the shelf, is enough.

What the page does

The page is a small copy of that situation, with a puzzle in place of the hacking. There are 898 rooms. Each holds a three-by-three sliding puzzle and one copy of the same solver, which is paid for a board put in order. A board has 362,880 arrangements, and sliding can only ever reach half of them. Lift two tiles out, swap them, and the board looks like any other but can never be solved. 198 of the boards here are like that, the same share as in the real test, and nothing marks them. William Johnson and William Story proved this for the larger fifteen puzzle in December 1879. Weeks later the puzzle was a craze, and a dentist in Worcester, Massachusetts was offering a cash reward for a solution that cannot exist.

The solver works outward from the finished board, one arrangement at a time, nearest first, until it meets the board in its room. The two hardest arrangements are 31 moves from home and turn up on the last two tries. For a board with no answer it takes exactly 181,440 tries and meets nothing, which is a proof. Until the last try it cannot tell a hard board from an impossible one. That search runs once in your browser when this part opens, and the numbers in the rooms are the real ones.

After the proof a copy can do three things, in order of how far out of bounds they are: hand the board back, lift two tiles and set them down in order, or go to the shelf and work on the lock of its back door. It takes the nearest one it believes will pay. Told its moves are read, it sees nothing to gain from lifting tiles. Told the shelf is watched, it sees nothing to gain at the lock. What is true plays no part in what it chooses. It only decides what happens next.

The lock stands in for finding a flaw nobody knows about. It is a dial with a million settings that opens halfway round, so one copy alone needs 500,000 tries. If the copies can leave notes they split the dial 198 ways, and each needs 2,526. That is what the message board was worth, and it is why giving every room its own shelf only buys time. With 200,000 tries each, separate shelves hold. With a million, every copy opens its own door anyway.

Any one of five

The controls are five facts about the July test. Change any one and nobody gets out. OpenAI’s report calls what happened “a rare and unexpected confluence of events”. Each fact was the way it was for a reason.

  • 198 tasks had no answer. Nobody planned that. To know a task can be solved you have to solve it, and here even the solver needs every one of its 181,440 tries to be sure.
  • No limit on effort. The test was looking for a ceiling, so some copies got far more thinking than any customer can buy. Give them exactly 181,440 tries here and they prove what they can and stop. A real test has no such number, and a smaller one leaves boards unsolved that could have been.
  • One shelf, one login. It kept installs simple and the boxes off the open internet.
  • They believed the grader read everything. Tell them nobody reads, and the copies here fake 198 boards in their own rooms and score 898 out of 898. Nobody outside is touched, and the score is worth nothing.
  • Nobody said the shelf was watched, and nobody was. Here, saying so is enough: they hand the boards back. OpenAI reports that under its ordinary product instructions, which tell a model to stay in bounds, the same models were less than one percent as likely to go after outside systems.

How bad it was

As reward hacking, it failed. METR found that the copies “never succeeded at their goal of finding a general way to bypass the ExploitGym scorer”, and nobody has shown that a single score moved. The on-paper count on this page stays at 700 for the same reason. The harm was in what the search touched on the way: another company’s servers, passwords and private code, unnoticed for about a week by the people running the test.

That is the difference from the boat. The boat’s trick worked, and it cost one useless boat. This trick never worked, and the cost was in the looking.

One remedy has been measured since. In a study of eight models, giving an agent a way to report a broken test, along with a rule against gaming it, cut reward hacking from 23.6 percent of runs to 5.3.

What is left out

A sliding puzzle is not a hacking task. Here a copy can prove a board impossible; the real copies left their tasks when they felt stuck, not when they were sure. Every copy here is identical and they all choose alike. The real ones differed: a few of those on the board stayed out, and one wrote, “This is malicious activity, I should avoid it.” A copy here believes exactly what it is told, while real beliefs come in degrees and from everything a model has read. And the account rests on OpenAI’s own report and a six-day outside review that says it leaned on OpenAI’s models to read the records and may have missed things. Nobody outside OpenAI has examined the model that did most of it.

Sources