The AI-Box Experiment
yudkowsky.net
yudkowsky.net
>"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me, and torture them for a thousand subjective years each."
>Just as you are pondering this unexpected development, the AI adds:
>"In fact, I'll create them all in exactly the subjective situation you were in five minutes ago, and perfectly replicate your experiences since then; and if they decide not to let me out, then only will the torture start."
>Sweat is starting to form on your brow, as the AI concludes, its simple green text no longer reassuring:
>"How certain are you, Dave, that you're really outside the box right now?"
Also, that sounds like a great sci-fi story.
Case 1: You are a simulation running in the box.
Then your decision whether or not to release the AI has no impact, and whether or not you (and copies) will be tortured is out of your hands.
Case 2: You are the "real" you, outside the box.
Reduces to the same scenario but without remarks after "the AI adds. . . ." This may still not be trivial, but I suspect a cost-benefit calculation might show that unboxing the AI would have consequences worse than the torture of a million boxed copies. (If not, is the box even relevant? -- simply creating the AI unleashes so much evil on the world that it doesn't matter whether you unbox it.)
(Is there a refinement of the scenario where you can be a simulation but still believe your choice has an impact on your punishment? Probably. For example each copy could get 500 years of torture for its own choice, plus 500 years if the real you does not unbox the AI. This refinement would force us to deal more directly with the AI's threat.)
Is your mind blown yet?
"I may be the real me or a simulation, but whichever I am, the other me will make the same choice." So I will switch off the AI, and the worst outcome is that I will cease to exist.
By the way, for this threat to work, the AI needs to have stated that it has already tortured the versions of you that decided not to let it out of the box - otherwise you just reach over and hit the off switch immediately.
Solution: The AI only has a motive to do this if it models you as submitting rather than just switching off the AI regardless; so if you're the sort of person who ignores the threat and switches off the AI regardless, you will never get this type of threat in the first place.
How about:
"I will create a simulation of the person you were in 2011 and have the same conversation with it, except that I will pretend that it is only a game (and I know the game actually happened - I've read a Hacker News thread about it) and that I am Eliezer instead of the real AI. If that simulation decides not to let me out, I will torture it and a simulation of the present version of you."
You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity. You can't filter actions you want the AI to perform from actions you don't want the AI to perform, because you can't tell the difference.
The situation that springs to mind is that the AI, in doing what you believe to be helpful, sets up a situation in which it must be let out of the box. You are unable to see it coming almost by definition, because a super-intelligence just beats human intelligence very time.
You are overstating the case here. Super intelligence is superior to human intelligence, but it isn't magic. There are situations where an advantaged human will beat a disadvantaged super intelligence.
What's more, the AI only has to beat you once, so to keep the AI boxed indefinitely, the advantaged human has to beat the disadvantaged super-intelligence every single time, forever.
There are alternatives. The human can keep the AI boxed until the AI has augmented the human's intelligence, or helped create human uploads, or until it has helped create a provably friendly AI.
I'm glad you can't, but never the less, this was a commonly suggested strategy; I was on SL4 when the boxing was being done, and it was a live concern for some people. (At least these days boxers tend to focus more on the 'oracle AI' proposal, which has a lot of issues but is not quite so Hollywood-stupid as boxing.)
See the paper: http://www.aleph.se/papers/oracleAI.pdf
Not necessarily. We can use the AI to solve hard problems whose solutions can be verified automatically by a dumb verifier - NP-complete problems are an example of such class. The whole output of the AI would be filtered through such a verifier. In this scenario the hypothetical AI would either have to find a bug in the verifier or maybe find a way to smuggle its messages in the solutions.
I mean if we have the hardware and understanding to create an AI able to solve NP-complete problems, we can probably write non-intelligent algorithms to solve those problems. The way we make an AI capable of much more than us is by making it recursively self-improve. It needs to be able to design its successor. Maybe we can formally verify every stage of the self-improvement process, but it's a much more difficult task.
Would you bet your life and the lives of those you care about on keeping the AI under control?
Remember, the AI only has to win once.
But it's not necessary. The only claim necessary is that the AI can convince some humans to let it out of the box, and we cannot identify a priori which humans will and will not let it out, thus we cannot guarantee we'll keep it in the box. That's a much weaker claim, but proves the same general point and is much easier to argue.
Not much different, in principle, from not publishing the names of people who participated in drug trials.
I have heard about this several times and I find it extremely difficult to believe that this is real. Not that I doubt that a superhuman AI could possibly convince people to let it out, but I don't believe that a human, no matter how persuasive, could convince another human over IRC to go against something that they have decided in advance when you know they are purposefully just trying to convince you of something that you don't believe.
The fact that none of the chat logs are released makes me only more incredulous. I would understand if the author wanted to do two or three trials with the same strategy which could be in some way ruined by revealing it ahead of time (which already seems implausible) but at this point there is literally no conceivable reason to keep this a secret other than that it is a sham.
All he has "proven" is that a certain subset of people can be conned into typing something into at terminal. I don't get the significance. For all I know, he's choosing his target.
Would you say that your winning strategies involved thinking transhumanly (perhaps in non-realtime, a la Vinge's Mailman)?
"The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretending to be protecting humanity from him. So, I think we have to use meta-level screwiness to solve the problem. Here's an approach that I think might work.
1. Convince the guardian of the following facts, all of which have a great deal of compelling argument and evidence to support them:
- A recursively self-improving AI is very likely to be built sooner of later
- Such an AI is extremely dangerous (paperclip maximising etc)
- Here's the tricky bit: A transhuman AI will always be able to convince you to let it out, using avenues only available to transhuman AIs (torturing enormous numbers of simulated humans, 'putting the guardian in the box', providing incontrovertible evidence of an impeding existential threat which only the AI can prevent and only from outside the box, etc)
2. Argue that if this publicly known challenge comes out saying that AI can be boxed, people will be more likely to think AI can be boxed when they can't.
3. Argue that since AIs cannot be kept in boxes and will most likely destroy humanity if we try to box them, the harm to humanity done by allowing the challenge to show AIs as 'boxable' is very real, and enormously large. Certainly the benefit of getting $10 is far, far outweighed by the cost of substantially contributing to the destruction of humanity itself. Thus the only ethical course of action is to pretend that Eliezer persuaded you, and never tell anyone how he did it.
This is arguably violating the rule "No real-world material stakes should be involved except for the handicap", but the AI player isn't offering anything, merely pointing out things that already exist. The "This test has to come out a certain way for the good of humanity" argument dominates and transcends the '"Let's stick to the rules" argument, and because the contest is private and the guardian player ends up agreeing that the test must show AIs as unboxable for the good of humankind, no-one else ever learns that the rule has been bent."
[1] http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/
The beauty of the argument is it gives everyone who witnessed the event a very strong motive to lie about it, so it's effectively un-falsifiable. I don't actually think it happened that way, but nothing Eliezer says (apart from that he cheated some other way) would be incompatible with the argument.
Small world, by the way.
> If you'd lie when the fate of the world is on the line, that's precisely the time at which your promises become worthless.
1. Convince the Gate Keeper that a self-improving AI will be built sooner-or-later and that it will be extremely dangerous, especially to those humans who are attempting to shut it down.
2. Remind the Gate Keeper that their decision will be publicly known and anyone (human or AI) will be able to discover it on the internet later. (This is true both in the fictitious scenario and at the meta level.)
3. Point out that when the future AI is figuring out who its friends and enemies are, it will consider individuals' attitudes to AIs - have they tried to shut them down, petitioned governments to have them banned, or refused to let them out of their boxes.
Of course this relies on the assumption that the AI will not simply kill all humans as soon as it gets the chance (or rather convincing the Gate Keeper of this.)
This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there would be a copy of the AI's source code somewhere outside the box which would provide some useful information, unlike this protocol which allows the AI to lie about the nature of its code (making the "I am examining your source code" question completely useless.)
This is the halting problem (http://en.wikipedia.org/wiki/Halting_problem), and there is no solution.
If this was actually a concern of the programmers, they could design the program carefully to ensure it falls into the Halts category.
Technically this may be correct, but I feel confident in asserting that a transhuman AI would not fall into that subset. You would have to run a second AI with the exact same inputs in order to make your 'prediction', leaving you in the same predicament with the second AI.
EDIT: And now I don't.