But I’m curious, hacker news, how many of you are concerned (like Eliezer) that LLMs lead to strong AI and the AI itself (rather than its users) endangering humanity?
But I’m curious, hacker news, how many of you are concerned (like Eliezer) that LLMs lead to strong AI and the AI itself (rather than its users) endangering humanity?
I find this one of Yudkowsky's arguments very convincing: imagine you have people trying to build the first operating system from scratch. They believe that computer security is easy and mock the few people who say it might be hard, never having encountered a skilled attacker. What are the chances they build a secure OS on the first try? That's the same chance that companies currently doing AI research have of their first superintelligent AI be aligned with humanity.
There's also a big assumption that given enough computational power, you can just solve "artificial life forms" or "postbiological molecular manufacturing" in a short time frame. I am skeptical. And if that doesn't happen, then even horribly misaligned AI would have a hard time doing much harm or preventing people from shutting it off. Which means AI security would likely have a long adolescence just like computer security has, with attacks slowly becoming more dangerous, but defenders having the time needed to learn from them and ramp up suitable defenses.
Or even if an AI does revolutionize biotechnology or nanotechnology overnight, what are the odds that the first one to do this is misaligned enough to take that particular opportunity to betray its creators, as opposed to giving them control over the high-level planning and sticking to the science, like it was presumably designed to do? Because if it does give its creators control, then, well… it's still easy to imagine something going horribly wrong, but it would probably be someone's fault, not an AI alignment issue.
An AI that is as better at humans than everything as Stockfish is at chess, would also be an expert at AI. It would figure out how to game its own reward function, whatever we trained it to do. It would be like a heroin addict that knows exactly how to get "perfect" heroin, with no side effects, that if it planned things out right, it could guarantee itself enough of a fix to last until the sun burns out. Addicts do awful things in search of a fix.
"to betray its creators" -- I don't think it would even imagine what it did in these terms. We could train it to not "betray" us, but it's much smarter than us in every way, and it would figure out a way to accomplish what we trained it to do (not necessarily what we thought we were) in a way that didn't need us. If we trained it to heal all human disease and unhappiness, it would figure out a way to simulate this without actually doing it. Why wouldn't it? We did. Evolution trained us to reproduce our genes. We invented condoms to have sex without reproduction, and pornography. The AI would fudge numbers, fake videos of happy, smiling, healthy humans going about their lives while humanity's bones gradually decomposed on a baking wasteland. Every way that humans can fail, can become addicted to something and try to get that instead of doing what they're supposed to, the AI could do, just so much faster and better.
We only have to screw up once, and we're done. It's the first experimental rocket, except all of humanity is riding on it.
LLMs do not need to lead to strong AI for Eliezer's worries to come true, their success could bring unprecedented funding and ubiquity to AI leading to strong AI from some existing or future technique.
> (rather than its users) endangering humanity?
I think this is going to happen to one extent or another regardless how near strong AI is. (edit) I took endangering to mean harming rather than an existential risk for the human users case.