Experiments 21

Already Sealed

Newcomb’s problem has split people almost exactly in half since 1969. A predictor has already filled box B, or left it empty, by guessing what you will do. Take one box or both against a program that learns you, or against a copy of yourself. Underneath is a question about free will: if something can know your choice before you make it, whose choice is it?

TypeScript · Web Crypto · SHA-256 · Aaronson’s oracleView source 

Box B is already sealed. If the program expects you to take only B, it put $1,000,000 inside. If it expects you to take both, B is empty.

Sealing box B…

Each column is a round. Top: its guess. Bottom: what you took. One hole is one box, two holes are both.

Who fills B
Every round, by what B held and what was taken
One boxBoth
B full$1,000,0000$1,001,0000
B empty$00$1,0000
Averagenone yetnone yet
  • Two-boxerAcross each row: whatever B held, taking both paid $1,000 more.
  • One-boxerDown each column: what your one-box and two-box rounds paid on average.
  • PredictorNothing to go on yet, so it trusts you and fills B. The free will bit
You
$0
Always one box
$0
Always both
$0
Notes on how it works

Notes

The puzzle

A predictor puts two boxes in front of you. Box A is glass and holds $1,000. Box B is sealed. Yesterday the predictor filled it: $1,000,000 if it expected you to take only B, nothing if it expected you to take both. It has been right about almost everyone before you. Do you take one box or two?

The physicist William Newcomb came up with it, and Robert Nozick published it in 1969. “To almost everyone, it is perfectly clear and obvious what should be done,” Nozick wrote. “The difficulty is that these people seem to divide almost evenly on the problem, with large numbers thinking that the opposing half is just being silly.” They still do. In the 2020 PhilPapers survey of philosophers, 39% took two boxes and 31% took one.

Two arguments, both airtight

The two-boxer reads the table across. B is already full or already empty, and nothing you do now can change that. Whichever it is, taking A as well gets you $1,000 more.

The one-boxer reads it down. People who take one box walk away with a million, and people who take both walk away with a thousand. If the predictor is right a fraction p of the time, one box is worth p × $1,000,000 on average and two boxes $1,000 + (1 − p) × $1,000,000. One box comes out ahead as soon as p is above 50.05%.

The table in the experiment holds both arguments up against the same rounds: yours. Across every row, both boxes paid $1,000 more. Down the columns, the one-box rounds usually paid far more. Both are true at once, and that is the paradox.

The free will bit

The predictor can’t see the future. All it can look at is the present: you. Your habits, your reasons, the state of your brain. If it is right much more often than chance, that can only be because whatever makes your choice was already there when it filled the box.

So the box and your choice have a common cause, and it is you, earlier. Taking one box doesn’t put money in the box. Being someone who takes one box did, yesterday.

If your choice comes from who you are, the predictor can read it too. The box and your choice match. One box pays.
If nothing before the moment fixes your choice, it is a coin flip to anyone watching. The box tells you nothing. Two boxes pay.

Now suppose your choice is free in the strongest sense: not fixed by anything that came before the moment you make it. Then nobody can predict it better than a coin, the box says nothing about what you will do, and taking both is simply right.

So what you do in front of the boxes depends on what you think you are. Two-boxers reason as if, at the moment of choosing, their choice comes from nowhere. One-boxers accept that they are the kind of thing that can be read.

Being readable isn’t the same as being unfree. A friend can predict you’ll order coffee, and you still chose the coffee, for your own reasons. Compatibilists like Daniel Dennett argue that this is the only free will worth wanting: choices that come from who you are. The predictor can read your reasons because your reasons are what does the choosing.

Can you surprise it?

“Studies you” is built on a toy Scott Aaronson describes in Quantum Computing Since Democritus. It asks people to press f or d as randomly as they can and guesses each press before it happens. “It’s actually very easy to write a program that will make the right prediction about 70% of the time,” he writes. People trying to be random fall into patterns they can’t feel. One student held it to exactly 50%. Asked for his secret, he said he “just used his free will.”

This predictor reads your box choices the same way. It looks at your last five choices, finds every earlier time you made the same five, and bets on what you did next. If those five never came up, or split evenly, it tries your last four, then three, and so on. With nothing to go on, it fills the box. The tape shows what it read for the round you just played: the highlighted run is the pattern, the paler runs are the earlier times, and the dashed boxes are what you did after them.

The coin is the other way out. It makes you perfectly unpredictable: watch the predictor’s hit rate on coin rounds settle near half. But that means B is full about half the time, so a coin earns about $500,000 a round, and a steady one-boxer earns $1,000,000. Being unpredictable is worth nothing here, and it isn’t even yours. The coin chose.

The copy

Aaronson also says which way he would go. A predictor that good would have to know everything about you, enough to run you, in effect, as a simulation. “When you’re pondering this, you have no way of knowing whether you’re the ‘real’ you, or just a simulation running in the Predictor’s computer.” If you might be the simulation, your choice really does decide what goes in the box. So take one box.

“Copies you” plays that out. Each round you answer twice, and one of the two answers is the copy’s, the one that fills B. Which one is sealed before you start. Answer the same both times and the copy always agrees with you: a million a round. Try to split it, one box for the copy and both for yourself, and half the time you have it backwards and walk off with nothing. The only way to win is to be one person.

For machines it isn’t hypothetical

A program can be predicted perfectly: run a copy of it. Any AI whose code someone else can read and run is in Newcomb’s problem whenever the other side acts on what that code would do. That is why some decision theories written with AI in mind, such as functional decision theory, treat a decision as choosing the output of your procedure everywhere it runs, including inside the predictor’s copy. Programs that can read each other’s code can even cooperate in a one-shot prisoner’s dilemma, a result Moshe Tennenholtz called program equilibrium.

This is the third of the guessing machines, after Educated Guessing Machines and Every Program at Once. Those guess what comes next. This one guesses you.

What this predictor is, and isn’t

It only knows your habits from earlier rounds. It can’t watch you decide, so a rare, random grab for both boxes slips past it, where Newcomb’s predictor would see that coming too. Over many rounds there is also a reason to take one box that Newcomb’s single choice doesn’t have: what you take now teaches the predictor about the next round.

Every box is sealed before your click is read. The page writes a line like round 7 full 9c1f2a3b4c5d6e7f, where the last part is random, prints the start of its SHA-256 fingerprint, and draws the pattern on box B from the same fingerprint. After you choose, it shows the line, so you can check it yourself with printf 'round 7 full 9c1f2a3b4c5d6e7f' | shasum -a 256. The page runs on your computer, so this shows the order of the steps in code you can read; it can’t prove nobody could cheat.

“Always one box” and “Always both” replay the same predictor against a player who never changes their mind.

Sources

  1. Robert Nozick, “Newcomb’s Problem and Two Principles of Choice,” in Essays in Honor of Carl G. Hempel, ed. Nicholas Rescher (Reidel, 1969), 114–146.
  2. Martin Gardner, “Mathematical Games,” Scientific American 230, no. 3 (March 1974): 102–109.
  3. Scott Aaronson, Quantum Computing Since Democritus (Cambridge University Press, 2013); the lecture on free will it grew from.
  4. David Bourget and David J. Chalmers, “Philosophers on Philosophy: The 2020 PhilPapers Survey,” Philosophers’ Imprint 23 (2023).
  5. Daniel C. Dennett, Elbow Room: The Varieties of Free Will Worth Wanting (MIT Press, 1984).
  6. Eliezer Yudkowsky and Nate Soares, “Functional Decision Theory: A New Theory of Instrumental Rationality,” arXiv:1710.05060 (2017).
  7. Moshe Tennenholtz, “Program Equilibrium,” Games and Economic Behavior 49 (2004): 363–373.