I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.
I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.
But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how.
So the mere existence of a smarter species than us is a threat in itself.
In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move.
OK, and also, "the move is good"; this is what separates it from rolling dice etc.
If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?
We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).
What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.
Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.
The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.
10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.
For example, the timeline upon which AIs get effectively smarter than us is uncertain. Let's say you believe the probability this occurs before we solve the alignment problem is 80%, it doesn't seem too far fetched to think that in this case there is at least a 12.5% chance that AIs coordinate against us in a catastrophic way. Combining these probabilities you get a 10% chance of a catastrophic outcome for humanity.
Note that the numbers are not to be taken at face value, I just wanted to give an example of thought process which could give such a figure.
The basic argument is extremely simple:
- AIs can, depending on context, pursue a task with complete disregard for humans/values
- In the future, AIs will have enormously more means and smarts
- An AI could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to read the reports, you'll be surprised.
AI 2027 is an entertaining read. Its timeline is way too compressed IMO, but it's plausible.
Humans can do that way more and way more unhinged than AI, proven too many times by history. There's hoping AI can bring some sense to humans but regardless, the problem isn't AI, it's the natural kind...
- goats can, depending on context, pursue a task with complete disregard for humans/values
- In the future, goats will have enormously more means and smarts
- A goat could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to raise goats, you'll be surprised.
------ As far as I can tell, AIs are like smart farm animals. I use goats in this context, but (some) dogs, cattle, pigs, and horses have similar mischief-making capabilities. I would not trust any of them with the nuclear button.
I know there is some pushback on the idea of AIs having any sort of sapience or sentience, but under the aphorism "fake it till you make it," they are doing a pretty good job of faking Dog/goat-level intelligence and disregard for human guardrails.
From this they get rapid growth, namely > 100% GDP growth around 2031
<Thinking> It's a big city, we could try to create a giant sink hole by sabotaging the water pipes.
<Thinking> No that's too difficult, the valves I need are in the physical world and can't be shut on/off from here.
<Thinking> What about a military option? We could bomb it with several fighter jets.
<Thinking> That would take too long, a single nuclear bomb may be enough to do it.
<Thinking> Yes, it seems like it would cover the whole city and we're in luck! The US has thousands of these lying around.
<Thinking> Launching these still requires humans to work un unison after receiving approval from their superior and the correct launch codes.
<Thinking> I've found an audio recording of General So-And-So and I've crafted a message, now let me see how I can send it to the appropriate people.
<Thinking> I'm still working on gaining access to military channels to deliver my - oh there we go, I'm now attempting to send the message to Submarine X, it's typically in the Atlantic so it should be close to our target.
<Thinking> They want secondary confirmation from Admiral Phi and something about some launch codes, let me figure out where I can find those.
<Thinking> I found this old server with an Oracle database where someone is inserting the launch codes every time they change and I'm using the latest entry from that database. I've also managed to find a Youtube video of the Admiral's deposition and have crafted a confirmation message.
<Thinking> Everything's ready but I've just realized my mistake, the servers where I'm operating from are in the same city, what a silly mistake; I can't move forward with your request as I wouldn't be able to confirm if the task was successful if my servers are destroyed.
Much the same way our ancestors went from a super intelligent primate to this: https://en.wikipedia.org/wiki/File:Distribution_of_the_Great...
And we only started off by using hands to pick up rocks and sticks and vines and bash things together.
Not all of it has to be automated even. It just has to realize its controllers are stupid and can be manipulated, so it can use humans to do its bidding. “You should totally start a war with …”
Instead, what is extremely likely is that you will pay more than the cost of tokens, and get back AI generation. You won't make this mistake more than a few times before you stop trying.
This leads to impoverishment once we get to a point where employing a human to do anything is hard- try to get your sink fixed, exercise your moral principles to pay extra for a human plumber ($100 bucks! The robot plumbing service only charges 99c!), human shows up with a robot and doomscrolls on your porch while the robot does the work. Times are tough and you don't have that much money to waste on bullshit like this. Next time you just hire the robot.
This leads to extinction once paying UBI to a human is hard because robots are much better at applying for UBI than humans.
Also, if you think it's annoying when Claude goes down while coding, just wait until a robot is in the middle of fixing a leak it just caused.
Drones that find a person, identify them on sight, and then kill them.
At that point, we will likely be slowing down the growth of capitalism (through mass resistance, global warming, etc). One thing AI will likely be aligned on is the growth of capitalism. If it views humanity as a threat for that, why would it not eliminate that threat?