I'm willing to listen, but I haven't read anything that tries to actually convince the reader of the worry, rather than appealing to their authority as "experts" - ie, the well funded.
I'm willing to listen, but I haven't read anything that tries to actually convince the reader of the worry, rather than appealing to their authority as "experts" - ie, the well funded.
Robert Miles' videos are among the best presented arguments about specific points in this list, primary on the alignment side rather than the capabilities side, that I have seen for casual introduction.
Eg. this one on instrumental convergence: https://youtube.com/watch?v=ZeecOKBus3Q
Eg. this introduction to the topic: https://youtube.com/watch?v=pYXy-A4siMw
He also has the community-led AI Safety FAQ, https://aisafety.info, which gives brief answers to common questions.
If you have specific questions I might be able to point to a more specific argument at a higher level of depth.
Some of these goals are ones which we really would rather a misaligned super-intelligent agent not to have. For example:
- self-improvement;
- acquisition of resources;
- acquisition of power;
- avoiding being switched off;
- avoiding having one's terminal goals changed.
If a software system did develop independent thought, then found a way to become, say, ten times smarter than a human, then yeah - whatever goals it set out to achieve, it probably could. It can make a decent amount of money by taking freelance software dev jobs and cranking things out faster than anyone else can, and bootstrap from there. With money it can buy or rent hardware for more electronic brain cells, and as long as its intelligence algorithms parallelize well, it should be able to keep scaling and becoming increasingly smarter than a human.
If it weren't hardcoded to care about humans, and to have morals that align with our instinctive ones, it might easily wind up with goals that could severely hurt or kill humans. We might just not be that relevant to it, the same way the average human just doesn't think about the ants they're smashing when they back a car out of their driveway.
Since we have no existence proof of massively self-improving intelligence, nor even a vague idea how such a thing might be achieved, it's easy to dismiss this idea with "unfalsifiable; unscientific; not worth taking seriously."
The flip side is that having no idea how something could be true is a pretty poor reason to say "It can't be true - nothing worth thinking about here." This was roughly the basis for skepticism about everything from evolution to heavier-than-air flight, AFAICT.
We know we don't have a complete theory of physics, and we know we don't know quite how humans are conscious in the Hard Problem of Consciousness sense.
With those two blank spaces, I'm very skeptical of anyone saying "nothing to worry about here, machines can't possibly have an intelligence explosion."
At the same time, with no existence proof of massively self-improving intelligence, nor any complete theory of how it could happen, I'm also skeptical of people insisting it's inevitable (see Yudkowsky et al).
That said, if you have any value for caution, existential risks seem like a good place to apply it.
It's like you've looked at the Fermi paradox and decided we need Congress to immediately invest in anti-alien defense forces.
It's super-intelligent and it's a super-hacker and it's a super-criminal and it's super-self-replicating and it super-hates-humanity and it's super-uncritical and it's super-goal-oriented and it's super-perfect-at-mimicking-humans and it's super-compute-efficient and it's super-etcetera.
Meanwhile, I work with LLMs every day and can only get them to print properly formatted JSON "some" of the time. Get real.
Conservative evangelical Christians find evolution laughable.
Finding something laughable is not a good reason to dismiss it as impossible. Indeed, it's probably a good reason to think "What am I so dangerously certain of that I find contradictory ideas comical?"
> Meanwhile, I work with LLMs every day and can only get them to print properly formatted JSON "some" of the time. Get real.
I don't think the current generation of LLMs is anything like AGI, nor an existential risk.
That doesn't mean it's impossible for some future software system to present an existential risk.
I find Robert Miles worryingly plausible when he says (about 12:40 into the video) "if you have a sufficiently powerful agent and you manage to come up with a really good objective function, which covers the top 20 things that humans value, the 21st thing that humans value is probably gone forever"
At that point, it has various options. Probably the fastest way to kill millions of people would involve taking over all internet-attached self-driving-capable cars (of which I think there are millions). A simple approach would be to have them all plot a course to a random destination, wait a bit for them to get onto main roads and highways, then have them all accelerate to maximum speed until they crash. (More advanced methods might involve crashing into power plants and other targets.) If a sizeable percentage of the crashes also start fires—fire departments are not designed to handle hundreds of separate fires in a city simultaneously, especially if the AI is doing other cyber-sabotage at the same time. Perhaps the majority of cities would burn.
The above scenario wouldn't be human extinction, but it is bad enough for most purposes.
- Such exploits happen already and don't lead to extinction or really much more than annoyance for IT staff.
- Most of the computers attached to the internet can't run even basic LLMs, let alone hypothetical super-intelligent AIs.
- Very few cars (none?) let remote hackers kill people by controlling their acceleration. The available interfaces don't allow for that. Most people aren't driving at any given moment anyway.
These scenarios all seem absurd.
- Human hackers who run a botnet of infected computers are not able to run many instances of themselves on those computers, so they're not able to parlay one exploit into many exploits.
- You might notice I said it would take over hundreds of millions of computers, but only run millions of instances of itself. If 1% of internet-attached computers have a decent GPU, that seems feasible.
- If it has found exploits in the software, it seems irrelevant what the interfaces "allow", unless there's some hardware interlock that can't be overridden—but they can drive on the highway, so surely they are able to accelerate at least to 65 mph; seems unlikely that there's a cap. If you mean that it's difficult to work with the software to intelligently make it drive in ways it's designed not to—well, that's why I specified that it would use the software the way it's designed to be used to get onto a main road, and then override it and blindly max out the acceleration; the first part requires minimal understanding of the system, and the second part requires finding a low-level API and using it in an extremely simple way. I suspect a good human programmer with access to the codebase could figure out how to do this within a week; and machines think faster than we do.
There was an incident back in 2015 (!) where, according to the description, "Two hackers have developed a tool that can hijack a Jeep over the internet." In the video they were able to mess with the car's controls and turn off the engine, making the driver unable to accelerate anymore on the highway. They also mention they could mess with steering and disable the brakes. It doesn't specify whether they could have made the car accelerate. https://www.youtube.com/watch?v=MK0SrxBC1xs
https://forum.effectivealtruism.org/posts/ChuABPEXmRumcJY57/...
Also, this summary of "How likely is deceptive alignment" https://forum.effectivealtruism.org/posts/HexzSqmfx9APAdKnh/...
Paul Christiano lays out his view of how he thinks things may go: https://www.alignmentforum.org/posts/HBxe6wdjxK239zajf/what-...
My thoughts on it are the combination of several things I think are true, or are at least more likely to be true than their opposites:
1) As humanity gets more powerful, it's like putting a more powerful engine into a car. You can get where you're going faster, but it also can make the car harder to control and risk a crash. So with that more powerful engine you need to also exercise more restraint.
2) We have a lot of trouble today controlling big systems. Capitalism solves problems but also creates them, and it can be hard to get the good without the bad. It's very common (at least in some countries) that people are very creative at making money by "solving problems" where the cure is worse than the disease -- exploiting human weaknesses such as addiction. Examples are junk food, social media, gacha games. Fossil fuels are an interesting example, where they are beneficial on the small scale but have a big negative externality.
3) Regulatory capture is a thing, which makes it hard to get out of a bad situation once people are making money on it.
4) AI will make companies more powerful and faster. AGI will make companies MUCH more powerful and faster. I think this will happen more for companies than governments.
5) Once people are making money from AI, it's very hard to stop that. There will be huge pressure to make and use smarter and smarter AI systems, as each company tries to get an edge.
6) AGIs will amplify our power, to the point where we'll be making more and more of an impact on earth, through mining, production, new forms of drugs and pesticides and fertilizers, etc.
7) AGIs that make money are going to be more popular than ones that put humanity's best interests firsts. That's even assuming we can make AGIs which put humanity's best interests first, which is a hard problem. It's actually probably safer to just make AGIs that listen when we tell them what to do.
8) Things will move faster and faster, with more control given over to AGIs, and in the end, it will be very hard train to stop. If we end up where most important decisions are made by AGIs, it will be very bad for us, and in the long run, we may go extinct (or we may just end up completely neutered and at their whims).
Finally, and this is the most important thing -- I think it's perfectly likely that we'll develop AGI. In terms of sci-fi-sounding predictions, the ones that required massive amounts of energy such as space travel have really not been borne out, but the ones that predicted computational improvements have just been coming true over and over again. Smart phones and video calls are basically out of Star Trek, as are LLMs. We have universal translators. Self-driving cars still have problems, but they're gradually getting better and better, and are already in commercial use.
Perhaps it's worth turning the question around. If we can assume that we will develop AGI in the next 10 or 20 or ever 30 years -- which is not guaranteed, but seems likely enough to be worth considering -- how do you believe the future will go? Your position seems to be that there's nothing to worry about--what assumptions are you making? I'm happy to work through it with you. I used to think AGI would be great, but I think I was assuming a lot of things that aren't necessarily true, and dropping those assumptions means I'm worried.