Timeless Decision Theory [pdf]
intelligence.org
intelligence.org
* http://intelligence.org/files/ProblemClassDominance.pdf
* http://arxiv.org/abs/1401.5577
(The original TDT manuscript occurred via pleas for 'something, anything to be published' that were satisfied by cleaning up an unfinished book draft and publishing it as-was.)
2) Can Paul see that you're not very self-aware? If Paul can't predict you at all, why does he believe that he can?
3) One of the applications of logical decision theories (TDT/UDT/etc) is between advanced agents or computer-based agents that (a) are very smart and might well be able to predict each other better than humans can and/or (b) might be able to look at each other's source code.
4) One of the other major applications is for the reflective consistency of self-modifying agents. The sort of self-modifying agents we're concerned with aren't human and can look directly at their own code; we should expect advanced forms of such agents to be correspondingly self-aware.
5) Many existing analyses, e.g. of Nash equilibria, assume that you know enough about the other agent to believe that, e.g., they will also be looking for a Nash equilibrium.
6) With that said, if you assume sufficiently great ignorance to break or prevent exploitation of any logical correlations, even correlations of the sort involved in "I know I try to maximize my utility" or "This agent and I are mutually ignorant of each other, and therefore will gravitate to a Nash equilibrium if there exists a unique Nash equilibrium", then you can end up outside the applicable domain of logical decision theories. (This may well be the case for actual humans! We can't read each other's source code, couldn't understand the brains if we saw them, don't know which decision theory the other person actually believes in, etcetera.)
With a perfect predictor, the hitchhiker problem seems simpler. Even if the agent participating in the hitchhiker problem is intentionally not self-aware (e.g. it was constructed by an agent tasked with submitting an agent to participate in the hitchhiker problem), the predictor can predict the actual behavior of that agent even if the agent can't, even if the actual behavior of the agent would be to "honestly" say that it intends to pay in the city and not decide otherwise until it reaches the city. So a non-self-aware program that would refuse to pay in the city would still die stranded in the desert. (In other words, just because you're intentionally non-self-aware doesn't prevent other agents from being you-aware.)
In the imperfect predictor version, such as with facial microexpressions that indicate lying, "not self-aware" seems equivalent to "self-aware but trained to not show the facial microexpressions that indicate lying"; both are cheats that ignore the purpose of the problem, and the problem can be reformulated to eliminate both such cheats.
MIRI is strange. Five years ago, I was convinced they were very misguided or cranks, but I've been reading some of their logic stuff and much of it is pretty neat. I also attended a talk by Yudkowsky at MIT and it was also very good. I'm not sure what to think now.
If the person has always taken B, then the Predictor will probably guess that the person will take B again, and therefore place $1,000,000 in B. By always choosing B, the person "wins".
If the person doesn't always take B, then the Predictor probably can't guess what the person will take, and B will be empty. By taking A and B, the person can only get $1,000.
The mechanism of Trust is exactly this. Whenever you're faced with a question, lying (to your advantage) will always increase the odds of a better outcome (like taking both A and B). However, that trick doesn't scale, and you can't use it anymore once a lie is discovered, once trust is lost.
I confess I haven't read the whole 100-page paper, but I've scanned it and I'm having trouble finding an algorithm concrete enough to be expressed in code.
>(Incidentally, don’t imagine you can wiggle out of this by basing your decision on a coin flip! For suppose the Predictor predicts you’ll open only the first box with probability p. Then he’ll put the $1,000,000 in that box with the same probability p. So your expected payoff is 1,000,000p2 + 1,001,000p(1-p) + 1,000(1-p)2 = 1,000,000p + 1,000(1-p), and you’re stuck with the same paradox as before.)
Even if you have "free will" and your actions are semi-random, that's not different than flipping a coin to make your decisions.
But you can only perform B-AB if you manage to fool the predictor and especially you can not do B-AB in a deterministic universe without free will (let's ignore that you don't even have control over the first step in that scenario). B is definitely the right choice in the first step and then changing to AB is the best choice in the second step, but whether you can actually do this depends on your stance on free will and the like.
Since you have imperfect information about yourself (this neuron talked to this neuron by these rules) you don't attribute the cause of your own behavior to anything other than "I did that".
I think the example from 1.3 is much more clarifying than the one-box vs two-box problem, there's no "predictor" involved
The minimal way to think about choice is to require the existence of an agent that is capable of producing binary outputs. Each output signifies a decision the agent has made. The agent is treated as a black box and whether the agent is deterministic or not remains unspecified and irrelevant.
Another approach is to attempt to endow the agents with some sort of free will.
An interesting realization is that there is no way to know whether humans or anything else for that matter qualify for the second option. There may exist a deity, alien or Matrix root who already knows all our choices before we do.
The funny thing is: we couldn't even tell the difference.
Now, imagine that there was a computer program that makes the same choices you do in the same situations, and the Predictor can read that source code before filling the boxes.
That's the intuition pump I use for thinking about Newcomb-like problems.