An Opinionated Introduction to Newcombology
On Causal, Evidential, and Functional Decision Theory
I talk about Newcomb’s Problem and related issues a lot, and I’d really like to have a single post to link to in order to give the background. Here, I want to describe (i) what Newcomb’s Problem is, (ii) why it matters, and (iii) the basic position on these sorts of questions.
1. A super quick introduction to decision theory in general
In decision theory, we talk about what acts you should choose given what your preferences are. We only care about what your preferences have to be like in order to be consistent and (instrumentally) rational; questions of whether preferences or actions are right or wrong, reasonable or unreasonable (if you think that’s different from rationality), etc. are left to other areas of philosophy. Outcomes are ways the world could end up, meant to capture things the agent cares about. States are ways the world could be, meant to be independent of your action. Acts are often thought of as functions from states to outcomes; your act in conjunction with the state of the world determines what outcome you get. Additionally, decision theorists represent the goodness of outcomes with real numbers, called utility values or utilities. Your confidence that a proposition is true is represented by a number from 0 to 1, called a probability or credence.
In this post, I’ll use the following notational conventions:
A, B, … refer to acts
s, t, … refer to states
x, y, … refer to outcomes
u refers to the agent’s utility function, a function from outcomes to real numbers
P refers to the agent’s credence function, a function from propositions to [0, 1]
Orthodox decision theory holds that when you’re acting under uncertainty, if you’re rational, you’ll choose the act that maximizes expected utility, the sum of the utilities of the possible outcomes weighted by their probability. If the possible outcomes are given by x1, x2, …, xn and the probability that xi obtains given that you choose A is denoted by PA(xi), then the expected utility of A is given by
Example: say it might rain, and I want to go for a walk, and I have to decide whether to take an umbrella. There are four possible outcomes: I go for a sunny walk with no umbrella, I get a sunny walk while having to carry an umbrella, I go for a rainy walk with no umbrella and get soaked, and I go for a rainy walk with an umbrella. Let’s say these outcomes have respective utilities of 10, 8, -5, and 2. We can represent the situation with the following table. The columns represent states, the rows acts, and the entries the utilities of the consequent outcomes.
Say the probability of rain is 20%; what should you do? With the formula above, we get
Therefore, given the stipulations, not taking an umbrella would be slightly better.
Something that basically all decision theorists agree about is the Principle of (Strict) Dominance. If an act A yields a better outcome than B no matter how the world is (i.e. no matter what state is actual), then we say that A strictly dominates B. Informally, the Principle of Dominance says that, if A strictly dominates B, then it is irrational to choose B. For example, in our umbrella case, suppose you love getting soaked, so taking a walk in the rain with no umbrella is 20 utils. The payoff matrix is therefore:
As we can see, no matter whether it rains or not, not taking an umbrella yields a better outcome. Therefore, the agent would be irrational to take the umbrella, no matter the probability of rain.
2. Act-State Independence, CDT, and EDT
Here’s an argument that you should pick up smoking, assuming you’d enjoy it. Either you’ll get cancer or you won’t. If you’ll get cancer, getting to smoke while having cancer is better than having to refrain while having cancer. Similar for if you won’t get cancer. Here’s the payoff matrix:
Not smoking is clearly dominated by smoking, so you’d be irrational to refrain from smoking.1
Obviously, this argument is terrible. Informally, the reason it’s terrible is that it treats whether you get cancer or not as a fixed fact independent of your choice, whereas in reality whether you smoke makes a difference as to whether you get cancer. In other words, in order for dominance reasoning to apply, the states and acts must be independent of each other.
But what does it mean for states and acts to be “independent”? There are two major views. On the one hand, we might mean evidentially independent: in the umbrella case, whether I take an umbrella or not gives me no evidence about whether it will rain, whereas here, me smoking is evidence I’ll get cancer. On the other hand, we might mean causally independent: in the umbrella case, taking an umbrella isn’t going to cause it to be more or less likely that it rains, whereas here, smoking does cause a higher likelihood of cancer.
Here’s a technical point: we phrase states in terms of acts and outcomes, through the use of conditionals. Instead of calling “it rains” a state, we call it “The world is such that if I don’t take an umbrella, I’ll get wet.” Likewise, in the smoking example, a state might be “The world is such that if I smoke, I’ll get cancer.” Since these are conditionals, they’ll be independent of the antecedents (the acts), giving us act-state independence for free (though I’m simplifying a fair bit). The two views—causal and evidential—then become two views about what kinds of conditionals these are. For the evidentialist, the state should be “If I perform A, then I’ll get the outcome x,” written A → x. When we speak of the probability of this conditional, though, we speak of the probability of getting x given that I do A, which is written as P(x | A). For the causalist, the state should not be an indicative conditional, but rather subjunctive; it should talk about what were to happen if I performed a certain act.2 The state, that x would result were one to choose A, is denoted by A □→ x. When I write probabilities, I’ll denote the probability that x would obtain were one to choose A as P(x | do(A)) (taking the notation from Judea Pearl).
Therefore, we have two decision theories. Evidential Decision Theory (EDT) says we should compute the expected utility of an act A by
whereas Causal Decision Theory (CDT) says you should compute expected utility by
This all may seem scholastic—why care about whether the conditionals are indicative or subjunctive, anyway? It’d be more concrete to have a case where the two disagree about what’s rational to do. And such a case, my friends, is given by Newcomb’s Problem. I’m going to introduce a character named Omega, who will pop up a lot for the remainder of this post. Omega is really accurate when it comes to predicting human behavior, with an incredibly small error rate of, say, one in a trillion. You also have very good evidence of Omega’s reliability, as well as his honesty. Moreover, we assume your choices won’t affect anything beyond your immediate outcomes; Omega won’t be sad if you choose a certain way (he’s just doing the experiments for fun), and your choice will have no impact on your rewards in future scenarios. Now, here is a case:
Newcomb’s Problem: Before you are Box A and Box B. Box B is transparent, and clearly contains $1,000. Box A is opaque, and contains either $0 or $1,000,000. You will be allowed two options: you can take only Box A (one-box), or you can take Box A as well as Box B (two-box). Whether Box A contains the million depends on Omega’s prediction about what you’d choose. If Omega predicted you’d one-box, he put the million in Box A. If Omega predicted you’d two-box, he put nothing in Box A. Omega already made his prediction and put the money in Box A or not; your decision has no causal effects on what’s in Box A. Do you one-box or two-box?
Here is a table representing the outcomes:
EDT says you should one-box, since you deciding to one-box is near-certain evidence that Omega put the million in Box A, since Omega is very accurate. If you one-box, you’ll almost certainly get a million; if you two-box, you’ll almost certainly get a mere thousand. Therefore, EDT recommends one-boxing. In contrast, CDT recommends two-boxing. You’ll take Box A no matter what, and your choice won’t affect how much money you get from that; all that your choice affects is whether you’ll get the extra thousand. So, for the CDTer, two-boxing is a dominating choice.
For the EDTer, the above table is a misrepresentation of the scenario, like my argument for why you should smoke. The states and acts aren’t evidentially independent. Instead, assuming Omega is equally good at predicting one-boxers and two-boxers, the EDTer will want
Sure, we don’t have a dominating choice, but given that Omega is super reliable, you’re overwhelmingly likely to end up in the left column, so you should ignore the right column.
In favor of two-boxing, we might argue as follows. Box A has $x in it; you can’t affect the value of x, the money is already in there or not. If you one-box, you get $x. If you two-box, you get $(x+1000). The latter quantity is always bigger than the former. Therefore, you should two-box. Additionally, imagine you had a friend who was allowed to peek into Box A, but who was not allowed to tell you what she saw. No matter what she saw, wouldn’t she tell you, with utter confidence, that you should definitely two-box? And if your friend would be right no matter what she sees, why not two-box to begin with?3
In favor of one-boxing, there is an extremely popular argument, which goes as follows: “If you’re so smart, why ain’cha rich?” Imagine your two trillion closest friends had their try at Newcomb’s Problem. Of your trillion two-boxer friends, all but one lucky individual ended up with $1,000. Of your trillion one-boxer friends, all but one unlucky individual are millionaires. Forget these fancy philosophical arguments about causation—wouldn’t you rather be a random person in the second group instead of the first? Don’t you want to be a millionaire?
In reply, the two-boxer admits that one-boxers are better off, but alleges that the choice is rigged—it punishes agents for being rational. Two-boxers are presented with a worse choice to begin with, simply because the predictor was able to discern that they’d make the rational choice. The one-boxer replies that this sounds like cope; if your choice predictably leads to you getting punished, then it wasn’t rational to begin with.
EDT basically tells you to pick the action with the best “news-value”—the action you’d be happy to learn was the one you’d taken. If you learned you’d one-box, you’d be very happy, because you’d probably be a millionaire. Therefore, EDT says, you should one-box. A major argument against EDT is that this way of reasoning sounds irrational, like wishful thinking—I should want to act in a way that causes the best outcome, not the way that merely correlates with getting a good outcome. Imagine that I felt sick, and considered going to the doctor. I reason: “People who go to the doctor are way more likely to have a horrible disease; therefore, I should not go.”
In reply, the EDTer will say this example doesn’t take all the evidence into account. If I feel some symptoms, my symptoms already give me all the relevant evidence of whether I have a horrible disease or not, and my choice of whether to go to the doctor on the basis of decision-theoretic considerations provides no additional evidence. Thus, I should go to the doctor.
In reply, CDTers give a more troubling case for EDT. Consider
Smoking Lesion: Scientists have just made a groundbreaking discovery—there is, in fact, no causal link between smoking and cancer. The reason smokers get cancer more often is because there is a genetic condition which makes you more likely to get cancer, and which makes you more likely to enjoy smoking. But everyone either has the genetic condition or not. Choosing to smoke will not cause a higher or lower propensity for cancer.
You don’t know if you have the genetic condition, and you can’t find out. You think smoking would be enjoyable, but most of all you don’t want cancer. Do you take up smoking?
CDT recommends smoking, which is intuitively correct—all smoking will do is add enjoyment to your life, in this example. Supposedly, EDT recommends refraining from smoking, since smoking would be evidence you have the condition and thus evidence you’ll get cancer. However, many proponents of EDT question whether the theory gives this recommendation. Like in the doctor case, the EDTer says that, realistically, you’ll be able to tell whether you have a stronger desire to smoke before you smoke. You’ll feel a “tickle,” an urge to smoke, and that will tell you everything there is for you to know about whether you have the genetic lesion, and your ultimate choice doesn’t provide any additional evidence. Many EDTers think Smoking Lesion is just like the doctor case above. This important counterargument is known as the “tickle defense” of EDT.
But what if, as is psychologically realistic, it’s not introspectively obvious whether you have a desire to smoke? Or what if everyone wants to smoke, and those with the genetic lesion just have a slightly stronger desire that pushes them over the threshold?4 It seems like the “tickle” needn’t be strong enough that your ultimate choice will give you no additional evidence that you have cancer. On behalf of EDT, I think we can reply as follows. Realistically, the lesion will cause you to smoke by way of your inclinations, even introspectively non-obvious ones. Thus, if I decide to smoke or not on the basis of abstract decision theory reasons, then that doesn’t give me any evidence as to whether I’ll get cancer. Most people aren’t thinking like this—they’re just smoking based on whether they feel like it. It’s not like refraining from smoking because it correlates with cancer correlates with no cancer. If, alternatively, we do assume the lesion correlates with how you think about decision theory—so, the lesion causes you to be more likely to refrain on the basis that doing so is evidence of not getting cancer—then this just becomes a version of Newcomb’s problem where Omega is replaced by a gene, and it no longer becomes so counterintuitive to refrain from smoking.
Now, you may be wondering: why care about Newcomb’s problem? Super-predictors like Omega don’t exist, and they probably won’t for a while. Why care about the difference between EDT and CDT if it only makes a difference in weird scifi cases? In reply, I think in fact Newcomblike problems are fairly common. Consider the still-unrealistic case:
Twin Prisoners’ Dilemma: You’re to play a game with a near-exact psychological duplicate of yourself. In whatever scenario, you two almost always choose the same way. Here’s the game: you both can either cooperate (C) or defect (D). If you both cooperate, you both get $1,000,000. If one of you cooperates and the other defects, then the cooperator gets nothing while the defector gets $1,001,000. If you both defect, you both get a mere $1,000. Neither of you know what the other chose until the game concludes.
Here’s the table:
Notice that this problem is nearly identical to Newcomb’s Problem, with Omega replaced by a twin. The twin simply acts as an accurate predictor of your choice. So EDT tells you to pick C, and CDT tells you to pick D.
Now, a question: how correlated does your partner’s choice have to be with your own in order for EDT to recommend cooperating?5 Let’s say that your credence that your partner cooperates/defects given that you cooperate/defect is p. Let’s also assume the dollars represent utility values. Then we have
To find the values of p for which EDT recommends cooperating, we write
This is interesting: the correlation between your choice and your partner’s barely has to be better than chance in order for EDT to recommend cooperating. I’d expect that, in real life, there would be such correlation between me and a random person—meaning that if I cooperate, I’ll be a little more confident they did too, and if I defect, I’ll be a little more confident they did too.
Human psychology is not transparent to us, and we learn a little bit about how others would act by observing how we ourselves act. Other people in the world act as slightly-accurate predictors of our own behavior, and therefore, Newcomblike cases are actually quite common! Remember, for the rest of this post, that Omega is really just a stand-in for the not-so-accurate predictors we encounter in our everyday lives, just stipulated to be nearly perfectly accurate for the sake of simplicity.
You’ll see more of the dialectic when I go through all the common Newcomblike cases. For now, let’s discuss an alternative to CDT and EDT.
3. Functional Decision Theory
Consider
Transparent Newcomb (just like Newcomb’s Problem, but you can see what’s in Box A): Before you are Box A and Box B. Box B is transparent, and clearly contains $1,000. Box A is also transparent, and contains either $0 or $1,000,000 (you can tell which). You will be allowed two options: you can take only Box A (one-box), or you can take Box A as well as Box B (two-box). Whether Box A contains the million depends on Omega’s prediction about what you’d choose. If Omega predicted you’d one-box no matter what you see in Box A, he put the million in Box A. If Omega predicted you’d two-box in either case (seeing nothing in Box A vs. seeing the million), he put nothing in Box A. Omega already made his prediction and put the money in Box A or not; your decision has no causal effects on what’s in Box A. Do you one-box or two-box?
Here, the CDTer two-boxes, for the same reason as before. But now the EDTer two-boxes as well, since they already know what’s in Box A, and their action provides no additional evidence as to whether they get the million or not. Say the EDTer sees nothing in Box A; if they one-box, they’re certain to get nothing, but if they two-box, they’re certain to get a thousand. Similar for if they see the million. Therefore, the EDTer two-boxes in Transparent Newcomb.
But, we may observe, it seems like the same considerations that weigh in favor of one-boxing in Newcomb’s Problem also favor one-boxing in Transparent Newcomb. My beloved EDTer, if you are so smart, then why ain’cha rich? Imagine your two-trillion closest friends again, one trillion of whom two-box in Transparent Newcomb, and one trillion of whom one-box. The latter are millionaires, and the former are thousandaires. Certainly, I want to be a millionaire, so why not join the one-boxers? (In my judgment, almost every argument for EDT over CDT functions as an argument for FDT—which recommends one-boxing in Transparent Newcomb—over EDT.)
Is there a decision theory that tells you to one-box in Transparent Newcomb, indeed one that tells you to always adopt the policy for action that makes you best off on average? One such view is called Functional Decision Theory (FDT). This view has a weird history; whereas CDT and EDT were salient in academic decision theory since discussion about Newcomb’s Problem began, FDT was formulated basically entirely outside of academia, mostly by an autodidact named Eliezer Yudkowsky. Yudkowsky is well-known for his work which gives personal advice on how to become more rational in your everyday life (see his Sequences as well as his rationality-focused fanfiction Harry Potter and the Methods of Rationality). From this work, he founded a community of rationality-enthusiasts calling themselves “rationalists,” centered on the website LessWrong. What became FDT is a result of a long back-and-forth between Yudkowsky and various people on LessWrong, especially Wei Dai.
Informally, FDT tells you to act on principles such that, on average, agents who have those principles do better in the relevant decision problems. Also informally, FDT tells you to act the way you would wish to commit yourself to acting. Any agent would want to force themselves to one-box before the prediction is made, in either the transparent or standard Newcomb Problem. Therefore, FDT says, one-boxing in either case is rational to begin with.
Less informally, FDT tells you to reason as follows. In the words of Yudkowsky & Soares, “Functional decision theorists hold that the normative principle for action is to treat one's decision as the output of a fixed mathematical function that answers the question, ‘Which output of this very function would yield the best outcome?’”6 In Newcomb’s Problem, the FDTer will reason as follows: “Suppose that FDT recommends two-boxing. Then Omega will have predicted I will two-box, and I will then two box, yielding $1,000. Now suppose FDT recommends one-boxing. Then Omega will have predicted that I will one-box, and I will one-box, yielding $1,000,000. The second case is better. Therefore, FDT says I should one-box.”
In other words, FDT says that the rational thing to do is whatever you would want to be the rational thing to do.7 Formally, FDT computes expected utilities as follows:
For each outcome, FDT has you imagine how likely that outcome would have been if FDT had recommended a given action.
The reasoning from Newcomb’s Problem applies exactly the same in Transparent Newcomb, where the FDTer will one-box and become a millionaire, whereas the CDTer and EDTer only get a thousand.
I should stress that this reasoning is kinda weird; FDT in fact recommends one-boxing, and indeed this is a necessary mathematical truth, but in order to figure it out, you have to imagine an impossible world were FDT actually recommends two-boxing. For some, reliance on counterpossible reasoning is a major consideration against FDT. In my opinion, though, this isn’t a good reason to reject the theory. Counterpossibles such as “If there were a counterexample to the Goldbach conjecture below 1010, this program would have output it” are manifestly significant, and figuring out how they work is everyone’s problem.8 Were someone to come up with CDT in relatively informal terms before the philosophy and statistics of causation were as developed as they are, this would not count against their theory. The goal of philosophy is to be correct, and you should expect that, sometimes, adopting a correct view will mean you have to sit with questions you don’t know how to answer yet.
Nevertheless, it is worth stressing that the FDTer still has a large check to cash. The formal issues involved with the kinds of counterpossibles FDT requires are very deep, and indeed spawn other problems. For example, you get a sort of reference-class problem; it’s easy enough for the FDTer to say you should cooperate in the Twin Prisoners’ Dilemma, since by hypothesis both parties are running FDT (assuming you run FDT). But what if your partner runs a slightly different algorithm? A moderately different one? (For the record, I think this is also everyone’s problem, at least insofar as you think mathematical truths can be explanatorily prior to people’s knowledge of them and their consequent actions; if you agree, then there are facts about which actions depend on which mathematical facts, so that we can evaluate the relevant subjunctives.)9
Functional Decision Theory is not wholly unprecedented in academia. Indeed, many of its core commitments were defended by the 20th-century philosopher David Gauthier. Additionally, the view bears a striking similarity to Kant’s Formula of Universal Law, which tells one to act only on a maxim (principle) that one can will to be a universal law (i.e. that one would decide should be the correct principle). Fortunately, while it went unnoticed for a while, FDT is gaining more discussion among academic decision theorists.
4. Various common Newcomblike Cases
Here, I will simply give a bunch of common decision theory cases, give our three theories’ verdicts on them, and maybe give a brief commentary. We’ve already discussed Newcomb’s Problem, Transparent Newcomb, Smoking Lesion, and Twin Prisoners’ Dilemma above. In all cases, again, we assume Omega (or whoever the predictor is) is very accurate and honest, you have good evidence of all that, and you only care about the specific outcome (e.g. money amounts, not dying, etc.). We also assume you have no way of randomizing your choice, except in a way that predictor can predict.
Death in Damascus: You live in Damascus. Death comes to your door and says “I am coming for you tomorrow.” In response, you can either stay in Damascus (D) or flee to Aleppo (A), though the latter will cost 1 util. Death is very good at predicting your choice, and if he’s in the same city as you tomorrow, you’ll die (-100 utils). However, Death has already made his prediction, and your choice won’t cause him to go to one city or the other.
EDT and FDT both say you should stay in Damascus; Death will almost certainly kill you no matter what, so at least you save the 1 util.
What CDT says about this case, if anything, is complicated. The tricky thing is that you’re uncertain about the causal structure of the case; if Death predicts A, then fleeing will cause you to die and staying will cause you to live, whereas it’s the reverse if Death predicts D. So, as the CDTer is getting on his bus to Aleppo, he says “Wait! Death will almost certainly be in Aleppo! Therefore, D will cause me to live.” But as the CDTer gets ready to stay in Damascus, he says “Wait! Death will almost certainly come here! Therefore, A causes me to live.” And so on. This is a feature of CDT known as instability.
In my opinion, the best response to this case for the CDTer was given by James Joyce (the decision theorist), who recommended adopting a random strategy, fleeing to Aleppo with a probability p. This means you have credence p that fleeing will cause you to die, and credence 1-p that fleeing will cause you to live. We can solve for the value of p that makes the choice stable. We have:
Setting these equal, we get p = 99/200, i.e. the CDTer wants to flee with a chance a little bit below 50%. Note that this strategy yields a worse outcome than simply staying in Damascus every time, because sometimes the CDTer unnecessarily pays the 1 util to flee before dying (since Death can still predict his choice).
For more discussion of this case and some variants, see “Cheating Death in Damascus” by Levinstein & Soares (link to pdf).
The Frustrater: You’ll be allowed to pick one box from among A, B, and C. C definitely has 40 utils in it. A and B together have 100, but the details depend on Omega’s prediction. If he thought you’d pick C, he put 50 utils in each of A and B. If he thought you’d pick A, he put 0 in A and 100 in B. If he thought you’d pick B, he put 100 in A and 0 in B.
This case and the ensuing discussion are from Jack Spencer (link to pdf). This case and the next one based on it are, I think, the most damning cases against CDT. Here, for the usual reasons, EDT and FDT both recommend taking C, for if you take A or B, you’re near-certain to get nothing. Between the choice of A and B, CDT will be unstable, so the CDTer probably adopt a 50/50 chance between them if they follow Joyce. However, one thing is for certain: CDT will deem it irrational to take C. This is because the two boxes add up to 100 utils, and therefore (because the contents of the box are causally independent of your choice)
If these two quantities add to 100, one of them has to be bigger than 40; two numbers less than or equal to 40 can only add up to at most 80. Therefore, whatever you think, CDT will certainly recommend taking Box A or Box B over Box C, whereas the intuitively correct choice is to pick C.
Now consider (still from Spencer):
Two Rooms: You can go to Room 1 or Room 2. Choosing Room 1 just gets you a guaranteed 35 utils. If you choose Room 2, you’ll play The Frustrater. Note that Omega’s prediction is made after you choose Room 2, if you do choose that.10
EDT and FDT recommend choosing Room 2 and getting the 40 utils. However, the CDTer will anticipate that, if they play The Frustrater, they’ll get nothing. Right now it’s before the prediction, so choosing 2 will cause them to get nothing. Therefore, the CDTer picks the 35 utils, when they literally could have just walked into the other room and grabbed 40!
XOR Blackmail: You’re 50% sure you have termites, which would suck a lot, amounting to -100,000 utils. Omega sends you an email saying: “Hi, I checked whether you have termites, and also predicted whether you’d pay me in response to this email (-10 utils). As it turns out, either you have termites or I predicted you’d pay me, but not both.” Do you pay up?
CDT obviously says not to pay up, since it makes no difference to whether you have termites.
FDT also recommends you not pay up. Should FDT recommend paying, then you still have a 50% chance of termites, and you’ll only get the email in the event that you do not have termites. Therefore, the EU is 0.5*-100,000 + 0.5*-10 = -50,005. Should FDT recommend refusing to pay, then you’ll only get the email in the event that you do have termites, but you won’t pay anyway. So the EU of not paying is 0.5*-100,000 + 0.5*0 = -50,000. So not paying is preferable.
It would seem that EDT recommends paying up. Omega is probably correct, so if you pay up, Omega probably predicted that, in which case, you have no termites. If you don’t pay, then Omega probably predicted that, and you probably have termites. Termites are really bad compared to paying up. Therefore, EDT says to pay up.
A prominent EDTer named Arif Ahmed has briefly argued that EDT doesn’t recommend paying up. You can see my reply to his argument here. I’m pretty confident he’s straightforwardly mistaken; however, Ahmed is an extremely good philosopher, and so it might turn out I’m in the wrong.
This case, in my opinion, is pretty damning for EDT. If EDTers pay up, they are highly exploitable; all Omega has to do is find some piece of potential bad news for you that doesn’t in fact obtain, and then he can blackmail you into paying him.
The key feature of this case is that, while paying is correlated in this instance with not having termites, people who are willing to pay are in the long run worse off than people who refuse. Having a disposition to pay doesn’t decrease your propensity to termites; all it does is determine whether you’d get an email when you have termites vs. when you don’t have termites.
Another problem for EDT is that EDTers will pay to have information withheld from them. See:
Newcomb and the Scoundrel: Omega has predicted whether you’ll one-box no matter what you observe. If you would, he put $1,000,000 in Box A. Otherwise, he put nothing in Box A. Box B is transparent and clearly has $10,000 in it. However, the Scoundrel has peeked into Box A, and he will tell you what’s in there unless you pay him $500,000. Omega predicted whether you’d pay, and what your choice would be consequent on paying or not.
Essentially, the Scoundrel is threatening to turn this into Transparent Newcomb. The CDTer will refuse to pay, since he’ll two-box anyway, so he leaves with $1,000. The FDTer will refuse to pay, since he’ll one-box either way, so he leaves the $1,000 behind and walks away with $1,000,000. The EDTer, however, one-boxes in Newcomb’s Problem and two-boxes in Transparent Newcomb. He anticipates that, if he doesn’t learn what’s in Box A, he’ll get the million, whereas if he does learn, he’ll two-box, and Omega will have predicted that, leaving him with $1,000. This is a difference of $999,000, so the EDTer pays the Scoundrel to stay silent. The EDTer then leaves with a mere $500,000.
Generally, learning things should be good (or at least not bad), unless the learning itself is bad (e.g. it causes psychological distress). What seems irrational is to pay not to learn things when the only consequence of learning is that you make an informed decision. A rational agent has no need to reject information for the sake of their future action; if the action you take under ignorance is so beneficial, then you can simply act that way after having received the information.
I should say that CDTers also face a similar problem, that of paying to do things they could just do for free. Suppose the Scoundrel charges $500,000 to force the CDTer to one-box, before the prediction is made. The CDTer will pay up. But why would a rational agent want to remove options from themselves? You can simply act the way you would pay others to force you to act, for free.
Here’s a case which bears some similarity to Transparent Newcomb:
Parfit’s Hitchhiker: You are dying in a desert. Dying is bad (-100 utils). Omega says he’ll save you if he predicts that, once you get back into town, you’ll go to the ATM and give him $100 (-1 util). If you manage to make it back into town, will you Pay or Refuse?
Assuming Omega predicts you’ll pay him, CDT and EDT recommend not paying. After all, you’re already saved, so paying is just a straightforward loss of utility that doesn’t cause any further benefit, nor is it evidence of any further benefit. Therefore, Omega will in all likelihood predict the EDTer and CDTer don’t pay, so they die in the desert.
FDT recommends paying up. If FDT recommends paying, you expect Omega to predict that, then you pay, and end up with -1 utils. If FDT recommends refusing, you expect to die, for -100 utils. Therefore, FDT recommends paying.
FDT’s recommendation seems insane to some people; you’re already in town, we stipulate that you only care about money and not dying, literally what reason do you have to pay? To lay my cards out, I don’t think there are any cases for FDT as devastating as Two Rooms is for CDT or XOR Blackmail is for EDT; for FDT, bad-looking actions are always offset by a higher expected reward. However, to those who think paying Omega when you get back into town would be crazy, that is probably the strongest consideration there is against FDT.
Here’s an even weirder case for FDT:
Counterfactual Mugging: Omega flipped a fair coin. If it lands tails, you’ll be allowed to give Omega $100 or not (-1 utils). If it lands heads, then Omega will give you $1,000,000 (1000 utils) if and only if he predicted you’d have paid the $100 if the coin had landed tails. If the coin lands tails, will you pay up?
When I teach decision theory, I’m surprised at how many students are receptive to transparent two-boxing and paying Parfit’s hitchhiker. However, Counterfactual Mugging is usually when they start to get off the boat. It’s clear that CDT and EDT recommend not paying. Additionally, FDT recommends paying, since FDT recommending paying gives agents a 50% chance of 1000 utils and a 50% chance of -1 utils.
I mentioned Gauthier before, a sort of proto-FDTer. Gauthier thought it was rational, for example, to pay in Parfit’s Hitchhiker; you simply form a commitment to pay because it saves your life, and then you pay. However, Gauthier would have thought you should not pay in Counterfactual Mugging. The reason is that Gauthier thought it was rational to abandon a commitment if following it would make you worse off than if you had made no commitment. In Parfit’s Hitchhiker, this condition is not met; sure, losing the 1 util and leaving me with -1 sucks, but if I had not committed myself to doing so, I’d be at -100 utils. In Counterfactual Mugging, though, paying does satisfy his criterion. The coin in fact lands tails; therefore, if you had made no commitment to pay, you’d still be at 0, whereas following through leaves you at -1. Thus, Omega will predict all this, and Gauthier ends up with nothing no matter what.
I take it this is the sort of reason why FDT’s verdict on Counterfactual Mugging is so counterintuitive to some; it asks you to make a sacrifice for a merely counterfactual benefit. At least in Parfit’s Hitchhiker, you’re reaping the benefits of your sacrifice; not so here. Nevertheless, EDTers and CDTers would be glad to pay someone 250 utils to force them to pay up; am I so irrational for simply causing the same effect for free?
5. Further Reading
For a general overview, see Arif Ahmed’s Cambridge Elements volume entitled Newcomb’s Problem.
The discussion of Newcomb’s Problem was kick-started by Robert Nozick. You can read his original piece in this pdf.
For an early but significant paper, see Gibbard and Harper, “Counterfactuals and Two Kinds of Expected Utility.”
For an early formal treatment of CDT, see Leonard Savage’s Foundations of Statistics.
For a defence and exposition of CDT, see James Joyce’s Foundations of Causal Decision Theory.
For a formal treatment of EDT, see Richard Jeffrey’s The Logic of Decision.
For a fantastic defence of EDT over CDT, see Arif Ahmed, Evidence, Decision, and Causality.
For an overview and defence of FDT, see Yudkowsky & Soares, “Functional Decision Theory.”
For a really good criticsm of FDT, see Wolfgang Schwartz’ post here
I originally got this way of introducing Newcomb’s Problem from Sarah Moss.
Selim Berker once noted in personal conversation that because of this, Causal Decision Theory should really be called Counterfactual Decision Theory, since causation and counterfactuals come apart.
I also got this argument from Sarah Moss.
If I’m not mistaken, I think I got this variant of the case from Harry Moss.
See David Lewis, Prisoners’ Dilemma is a Newcomb Problem.
This needs caveats—strictly, this slogan only applies in what Yudkowsky & Soares call “fair” decision problems, where your outcome depends on your behavior in that very decision problem. In a case where a predictor would punish you if he predicted you’d one-box in the standard Newcomb Problem, you would sure want two-boxing to be rational, but this fact does not impugn the rationality of one-boxing, since the punisher-related considerations are extrinsic to Newcomb’s Problem.
Even if you think counterpossibles are vacuous, we’re still doing something meaningful when we reason in terms of them. Perhaps they are a mere heuristic for something else; if so, then we should replace counterpossibles in FDT with whatever they’re a heuristic for.
It’s worthwhile to clarify what I mean when I say I believe FDT. Perhaps some of the formal problems are intractible, and an improved theory won’t make reference to counterpossibles. What I am confident of is that FDT gets the right verdicts in the cases, and that the verdicts are right for basically the reasons FDT says so. I think the correct theory will be at least as close to Yudkowsky’s formalism as different varieties of CDT are from each other.
It should be noted that this detail departs from Spencer, who has the prediction made before the choice of Room 1 and Room 2. If I haven’t made any mistakes, I think my setup is preferable, as it at least seems debatable whether the CDTer would choose Room 1 in Spencer’s version. Thanks so much to Hein de Haan for pointing this out in the comments.


"Note that Omega’s prediction is made after you choose Room 2, if you do choose that."
Perhaps it's worth noting that that's not true in Spencer's original problem statement, which (in note 6) says that the Frustrater made the prediction yesterday.
I wonder if you could cash out some of the FDT counterpossible reasoning stuff via some sort of relevance logic. You wouldn't need to actually adopt the relevance logic (or whatever) as "true," much less provide a semantics for counterpossibles, just base your calculations partially on what adopting this logic would imply. "If X is a theorem in some formal system I reject, then do A, otherwise do B," etc.