A Trilemma Regarding Newcomblike Problems
Another defense of Functional Decision Theory
If you don’t know what Newcomb’s Problem is and what Functional, Evidential, and Causal Decision Theory are, read this post of mine, or better, the paper on FDT by Yudkowsky & Soares
When you’re deciding what course of action is best, you need to reason about how things will turn out, supposing you act a certain way. Whichever act yields the best expected outcome supposing you perform that act is the act you should do. A suppositional decision theory is just a specification for how to determine what the world will be like supposing you perform a certain act. CDTers think we should imagine everything in the past stays fixed until your act, after which we look at the expected causal results of your act. EDTers think we should just imagine how you’d expect the world to be conditional on learning you performed an act; you should choose the action which, on your evidence, would correlate with getting the best outcomes. FDTers think we should imagine the world where FDT itself recommends that act, so that we can adopt the principle for action which will make the agent best-off in general.
A major sticking-point for many people on FDT is that it can recommend acting in a way that you know for certain will result in a very bad outcome. MacAskill gives basically the following case (modified by me, below is not a quote):
A highly accurate predictor has predicted how you will choose from among two boxes, boxes we call Left and Right. Right is empty, but you have to pay $100 to pick it. Left might have a bomb in it, and you’ll be told whether it does. The predictor predicted what you’d do conditional on finding out there’s a bomb in Left, and conditional on finding out there’s no bomb; if you’d pick Right in at least one of the two cases, he put a bomb in Left. Now you find out there’s a bomb in Left, and if you take it, you’ll burn to death.
Assuming the probability of an error on the predictor’s part is sufficiently small, the FDTer will want to pick Left, and then burn to death. Informally, the rationale is that people who adopt a principle of choosing Left no matter what are better off on average, since keeping $100 is worth a tiny chance of burning to death. As it happens, the FDTer in our thought experiment was just insanely unlucky.
Some people find this verdict alone to present a decisive case against FDT. I cannot empathize with this state of mind, I’ll admit. In the abstract, it wouldn’t be the craziest thing in the world to find out that performing an action that for certain has bad consequences for myself can be rational if it’s part of a plan that is advantageous to adopt; finding out that paying in Counterfactual Mugging is rational would not be the craziest thing ever, and I expect my 21-year-old-CDTer-self would have agreed had she known about Counterfactual Mugging. And MacAskill’s case contributes nothing new whatsoever to this matter; it just makes the outcomes more extreme, so as to (unintentionally, perhaps) manipulate the reader.1 It would be like if I said “What, you CDTer, you’d two-box even if the second box only had one cent in it, and the first box gave you eternal bliss if the predictor predicted you’d one-box?”
At any rate, intuitions aside, I think it can be compellingly argued that we should give up dominance—the principle that if an act A for certain causes a better outcome than the act B, then it is irrational to pick B.
First, some notation and terminology. I’ll denote outcomes by x, y, …, propositions by p, q, …, and acts by A, B, …. P will refer to the given agent’s credence function, and u to their utility function. I’ll write A ≺ B to mean that an act B is preferred to A, and similar for the other curvy ordering signs, and I’ll write A ~ B to mean that A and B are indifferent. A ≺p B means that, if the agent were to learn p, then they would prefer B to A. P(x | A) refers to the probability that the outcome x obtains conditional on the agent choosing A (the thing the EDTer cares about), and P(x | do(A)) refers to the probability that x would obtain were the agent to choose A (the thing the CDTer cares about). Finally, I’ll steal some machinery from Jack Spencer (link to pdf). An act A is said to “guarantee” an outcome x provided that, if the agent chooses A, they will have available to them a choice that yields x. An act A is said to “force” an outcome x provided that, if the agent chooses A, they will definitely get x as a result.
Here are five plausible principles:
Statewise Dominance: Let p1, p2, …, pn be a partition of probability-space (meaning that, for certain, exactly one of the propositions is true). Moreover, which of the pi is true is causally independent of the agent’s choice. Finally, suppose that, for each i, we have that A ≼pi B, and that there exists a j such that A ≺pj B. Then A ≺ B.
Guarantee Principle (due to Spencer): Suppose that A forces x, B guarantees y, and u(x) < u(y). Then A ≺ B.
Causal/Evidential Dominance: Suppose that P(x | A), P(x | do(A)), P(y | B), and P(y | do(B)) are all equal to 1, and that u(x) < u(y). Then A ≺ B.
Transitivity: If A ≺ B and B ≺ C, then A ≺ C.
Continuity: If u(x) < u(y) < u(z), then there exists a probability t such that a sure thing of y and a t/1-t chance of x/z are indifferent.
Informally, Statewise Dominance says that if you’d consider B a better choice than A no matter which pi you learned to be true, then you should consider B better than A before learning anything. The Guarantee Principle says that e.g. if one choice gives you a sure thing of winning $10, and another choice forces you to end up with $5, then you should consider the former choice better. Finally, Causal/Evidential Dominance says that if A is certain to yield x and B is certain to yield y (both in the causal and evidential senses), and y is better than x, then B is preferred to A.
Notice that from Statewise Dominance we can deduce a principle I’ll call Limited Stochastic Dominance, which states that if
p entails q
P(p) < P(q) (resp. ≤)
A yields y if p and x otherwise
B yields y if q and x otherwise
p and q are causally independent of whether one chooses A or B
and u(x) < u(y)
then A ≺ B. All this is just to say (in a specific case) that you should prefer (resp. not disprefer) actions that have a higher (resp. higher or equal) chance of a better reward. To see that Limited Stochastic Dominance follows, simply consider what A and B yield conditional on the three possibilities p ∧ q, ¬p ∧ q, and ¬p ∧ ¬q.
Unfortunately, though, given some very modest richness assumptions, only at most three of the principles 1-5 can be true. I assume Transitivity and Continuity are not up for grabs, so we have to figure out which of 1-3 to reject.2 Respectively, the first three principles are given up by EDT, CDT, and FDT.
This fact can easily be seen from a generalized version of Spencer’s “Frustrater” case, linked above. To build the case, I just need to suppose there are four outcomes x, y, z, and w such that u(y) > u(z) > u(w) > u(x) and such that a 50/50 gamble between x and y is preferred to z. (Mnemonic: x sux, y is grayt, z is zatisfying, and w is whatever.)
Now, suppose an agent satisfies Statewise Dominance, Causal/Evidential Dominance, Transitivity, and Continuity. Consider the following preliminary case:
The Frustrater: You will be allowed to choose one of three boxes, called Boxes 1, 2, and 3. Box 3 has a ticket for z in it, no matter what. Of the other two boxes, one of them has x and the other has y, though which is which depends on the predictor’s prediction. If the predictor thought you’d choose 3, he flipped a fair coin to decide which box gets you which outcome. Otherwise, if he thought you’d pick Box 1 or Box 2, then he put y in the box he thought you wouldn’t pick, while x is in the box he thought you would pick. All predictions and coin flips have been done, and what you do now has no effect on which box contains which reward/punishment.3
Here’s a table illustrating the outcome of each box conditional on each prediction/coin flip result:
What needs to be noted is that, if our agent can give any answer to this case at all, they will not pick Box 3. Let p be the proposition that Box 1 contains y, so that P(p) is the agent’s credence (at the time of choice) that Box 1 has y. If P(p) = 0.5, then by hypothesis, our agent will prefer the 50/50 risk between x and y to a sure thing of z. If P(p) > 0.5, then choosing Box 1 is preferable to an adequately-defined 50/50 gamble between x and y (by Limited Stochastic Dominance), which in turn is preferable to z (by hypothesis). So by Transitivity our agent will want to pick Box 1. Finally, if P(p) < 0.5, then our agent will similarly want to pick Box 2.4 The agent will, then, pick Box 1 or Box 2, and, because the predictor is super-accurate, they’ll almost certainly end up with x.
This verdict by itself is bad enough, but we don’t yet have a violation of the Guarantee Principle. To get that—and we’re still following Spencer here—we introduce the following meta-case5:
Two Rooms: You can enter Room 1 or Room 2. If you enter Room 1, you get w. If you enter Room 2, you get to play a round of the Frustrater. Importantly, if you opt for the latter, then the predictor will make his prediction after you choose Room 2 but before playing the Frustrater.
Choosing Room 1 will be preferable to a hypothetical choice that yields x for certain, by Causal/Evidential Dominance. Similarly, a hypothetical choice to get y for certain will be preferable to choosing Room 1. Briefly: EU(x) < EU(Room 1) < EU(y). Therefore, by Continuity, there exists a probability t such that choosing Room 1 is just as good as a gamble with a t chance of x and a 1-t chance of y. By Limited Stochastic Dominance, then, for any s < t, we have that choosing Room 1 is preferable to an s chance of x and a 1-s chance of y. We may suppose that the error rate of the predictor is s, which is less than t. But then choosing Room 2 is precisely the gamble mentioned, so our agent chooses Room 1.
In Spencer’s terms, Room 1 forces w while Room 2 guarantees z (since the latter gives you a later option that will certainly yield z). Moreover, z is better than w, so we have a violation of the Guarantee Principle. And choosing Room 1 is just crazy; why do that and get w when you can literally just walk into Room 2 and take z?
For those compelled by the Bomb case, we can specify our outcomes. Imagine that w represents taking the bomb and burning to death, that x represents 1 year of torment (and then dying), y represents n years of bliss for a sufficiently large n (and then dying), and that z represents getting a cookie. Any theory satisfying 1, 3, 4, and 5—including, therefore, Causal Decision Theory—will have you burn to death when you could have just gone to the other room and eaten a cookie instead (I believe there’s a Sartre quote about this). They will do this every time, by the way, whereas the FDTer in a generic version of MacAskill’s Bomb case will basically never burn to death, if the predictor really is accurate.6
So, we need to reject either Statewise Dominance or Causal/Evidential Dominance. I want us to reject the latter; but why not reject the former? To restate:
Statewise Dominance: Let p1, p2, …, pn be a partition of probability-space (meaning that, for certain, exactly one of the propositions is true). Moreover, which of the pi is true is causally independent of the agent’s choice. Finally, suppose that, for each i, we have that A ≼pi B, and that there exists a j such that A ≺pj B. Then A ≺ B.
Giving up Statewise Dominance doesn’t seem too crazy. After all, if you’re an EDTer, you’ll one-box in the standard Newcomb’s problem, whereas in Transparent Newcomb (where you can see whether the predictor predicted one-boxing or two-boxing), you’ll two-box. If p is the proposition that Box 1 contains the million, then for the EDTer, we have that Two-box ≺ One-box, even though One-box ≺p Two-box and One-box ≺¬p Two-box.
What this means for the EDTer is that they’ll pay you to not give them information. Ignorant of the prediction, the EDTer reasons that if they one-box they’ll probably get a million and if they two-box they’ll probably get a thousand, so they go for the former. Were they to learn the prediction, though, the agent’s choice would no longer be evidence that the first box has a million, so they’ll two-box. The EDTer can foresee that, if they learn the prediction, they’ll end up with $1,000; thus, they’ll pay up to $998,999 to someone who threatens to tell them the prediction, when they could have just decided to one-box despite knowing the prediction, thus getting a million.
Indeed, anyone who violates Statewise Dominance while obeying the other four rules (plus completeness and richness assumptions) can find themselves in such situations. Suppose the hypotheses of Statewise Dominance are met, so in particular we have A ≼pi B for all i and A ≺p1 B, but B ≼ A. We may assume (this is the richness assumption) that there’s a mild sweetening (slightly better version) of called A+ such that A+ ≺p1 B and B ≺ A+. The sweetening might be an additional $1, for example. Now, here’s the game. You don’t know which pi is the true one. For now, you have the option to be paid $1 to pick A over B. However, if you don’t pay up $t, you’ll be told which pi is true. In that case, you might learn p1 is true, in which case you’ll pick B. You want to avoid that; for sufficiently small t, you’d pay that much to avoid getting the information. (Note that we can only set this case up because the pi are not causally dependent on the choice, so that we don’t get weirdness surrounding the information having implications as to your future choice.)
Things get worse, namely, due to the case of XOR Blackmail, which I’ve written about here and here. Indeed, since writing those I have come up with an ingenious invention (if I’m not mistaken; no one has corrected me yet) called OR Blackmail. The specific case would go as follows:
OR Blackmail: You’re 50% sure you have termites, and if you do, that’d cost $100,000 in damages. You get an email from a super-accurate predictor saying: “I checked whether you have termites, and also whether you’d give me $1,000 in response to this email. I can verify that one or the other is true, or perhaps both. Were I to have found neither was true, I would not have sent this email.”
If you don’t pay up, you’ll be near-certain to have termites, but if you do pay up, you’ll be significantly less confident you have termites; probably no more than 50% confident, as I see no reason why getting the email and then paying up would be evidence for termites.7 Therefore, if the numbers are stipulated correctly, the EDTer will pay up.
For our Bomb fans, we can make the case more spicy:
OR Blackmail (Spicy Version): A super-accurate predictor scoured the history books, anthropological data, etc. to find some hitherto-unknown-to-you atrocity that killed a million people at some point in the past. You’re 50% confident that this particular atrocity occurred (also, the predictor sends these emails with fake atrocities to people who won’t pay up, so the email isn’t evidence the atrocity occurred). The predictor tells you that either the atrocity occurred and a million people died as a result, or you’ll blow yourself up and burn yourself to death with a bomb. By the way, you’re reasonably altruistic, and would gladly give your life if it diminished the likelihood of a million deaths by 50%.
Likewise, the EDTer will blow themselves up, here. And, likewise, they’ll do it every time.
It is not clear to me that this sort of argument can be made against anyone who accepts principles 2-5 while rejecting 1, but (i) EDT is the only well-motivated view that rejects 1 while also accepting 3, and (ii) paying to not receive information when you could just costlessly act as if you didn’t receive that information seems pretty bad on its own (you can figure out how to make this case into one where you blow yourself up every time).
Some people are going to try to appeal to resoluteness, here, which means adopting a plan and then sticking to it even though you’ll find a later choice disadvantageous. They’ll say the CDTer can just commit to choosing Room 2 and then choosing Box 3 in the Frustrater, and the EDTer can just commit to not paying the blackmailer, and then they can follow through on those commitments just because it was rational to adopt the plan. But that move just as well motivates resolute one-boxing in Transparent Newcomb, which means rejecting Causal/Evidential Dominance, which was supposed to be a premise in a powerful argument against FDT.
Suppose you were told that tomorrow you’ll face a round of Transparent Newcomb; if you can convince yourself to one-box, you’ll almost certainly be a millionaire, but if you can’t, you’ll almost certainly get just a thousand. (I am being nice and not making this about bombs.) And, alas, you have no way to force yourself to one-box, let alone trick the predictor, whether physically or by self-deception. You’re highly confident that you can only be a millionaire by actually deciding to one-box, and you can only get yourself to do that by deciding it’s the rational thing to do.
You could imagine having an argument with your future self. Your future self says either “Look, the million is in Box A already! I’d be insane not to take the extra thousand in Box B as well!” or “Look, the million isn’t in Box A anyway! I have no reason to not take the thousand!” But your current self says “I know, I know, but I really want you to take just Box A; if I can just freely choose to do that tomorrow, then I’ll almost certainly get a million, whereas I’ll only get a thousand otherwise.” Your future self then says “Too bad, you’re not the one choosing, I am, and I want the extra $1,000.”
But would it be the craziest thing in the world if it turned out that the perspective of your future self doesn’t automatically win? You are, after all, the same person at both times, and so you get to pick which perspective gets privileged. If we don’t beg the question of which to privilege, then there is a lot to be said in favor of privileging the earlier (namely, that it makes you better off in general), and little to be said in favor of privileging the later other than that it feels like the rational thing to do.
It’s no surprise that if we fail to take a time-neutral approach, where we privilege one perspective or the other in accordance with their general benefits, then we run into huge trouble when it comes to acting over time, as we saw with CDT and EDT. Agents who obey these decision theories treat their future selves like separate agents whose choices they have to plan around, and it’s no wonder such a constraint keeps them from getting what they want.
I should say, though, that even if I’ve convinced you to reject Causal/Evidential Dominance, that doesn’t mean you have to accept FDT. You could, for example, have Gauthier’s view, on which it is rational to abandon a commitment if following it would make you worse off than if you made no commitment at all (so Gauthier one-boxes in Transparent Newcomb, but pays the $100 in MacAskill’s case). But I hope to have removed a major barrier to finding FDT plausible.
There are also some subtle structural differences, but they don’t matter here.
Actually, Continuity is false in general, in particular when we’re dealing with infinite goods (e.g. eternity in Heaven). But we’re not dealing with anything like that here.
In Spencer’s original case, he was more concrete. Box 3 had $40 for certain and Box 1 and Box 2 had $100 between them. If the predictor predicts you go with 3, he put $50 in the first two. Otherwise, he put $100 in the box he thought you wouldn’t pick and $0 in the one he thought you would.
Cf. Spencer, p. 3.
Cf. Spencer, p. 4 (above is not a quote).
“But it’s a stipulation of the case that the predictor predicted you’d pick Right, so the FDTer does die every time in that case.” That is not how we should individuate decision problems; it would be like presenting the case, “You can take a bet where a fair coin is flipped, where you get $100 if heads and lose $10. By the way, unbeknownst to you, the coin will in fact land tails. Do you take the bet?” and then criticizing a decision theory that recommends taking the gamble on the basis that it loses “every time.”
We could stipulate more about the case to ensure this. Say the email is sent to everyone who has termites or would pay up, and that there is no correlation between having termites and being willing to pay up. Given that you paid up, then, you’ve gotten no new information about whether you have termites.


Really cool article
One point though: You are writing that the 1-5 agent would walk into room 1 against a low-accuracy frustrator, right? Shouldn't it be the opposite, against a low accuracy frustrator he chooses 2, and if the predictor is more accurate than t he chooses 1? Is this just a typo or am I missing something?
Flo have you engaged at all with Olé Peters’ work on non-ergodicity and path dependence and its implications for decision theory? He calls it “ergodicity economics” but I think this is a confused label since it’s specifically non-ergodicity which is being emphasized.
Anyway I’d be really curious how it struck you.