Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
I think the main argument for ASI is something like (i) extrapolating the progress from the past 10 years into the future, (ii) rapid progress apparently still being made, and (iii) seeing no obvious theoretical limitations.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful tend to see risks to their own status (in a brain stem primate status assessment sense) as globalized existential risks. E.g. Rome, as it fell, saw this as the literal end of the world.
Democratized AI could be a threat to both intellectual and financial privilege by making big ideas, science, and high-labor enterprise more achievable by everyone. It could be the white collar intellectual labor equivalent of the combine, the automated weaving machine, or... the crossbow. The intellectual equivalent of the crossbow would be automated fact checking at scale to defeat propaganda, a labor union using a superhuman AI to coordinate its organizing efforts using game theory, etc.
Hence the desire of the existing elite to make absolutely sure they control it. For our own good, of course.
> How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do. I haven't given this much thought, so maybe I'm missing something, but I don't see how this follows. One possible solution: to avoid overfitting, can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
Regarding the X-risk vs D-risk: I think how one weighs these risks partially depends on what one thinks the capabilities of the models are. Call me a boot-licker, but if the models get smart enough to explain, in detail, to any psychopath, how to construct a bomb or synthesize a deadly virus, I don't think benefits society to distribute them widely. Therefore, to argue for widespread distribution you have to argue that either (i) the models aren't that capable or (ii) the guardrails are robust enough to prevent them from being used in catastrophic ways by bad actors. I think we may be rapidly approaching a time where neither of these hold. Having said that, I certainly agree that the D-risk is also real.
> can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
That's the same as setting a fixed metric for intelligence, like IQ testing, and goal seeking that.
The problem is what happens when you max that out. How do you set the next metric? Now you're back to the recursive problem of your metric "begging the question." I don't think you can get smarter by seeking "I'm smarter because I think I'm smarter and I'm right because I'm smart."
On X-risk stuff:
They can already do what you say. Try asking an ablated 30B model on your laptop how to weaponize anthrax. The answer isn’t bad. It’ll tell you how to cook meth too.
I studied undergrad biology and walked out with the knowledge to create some damn evil things if I had the right lab, time, and no conscience. This was pre AI. The recipes for a lot of nasty stuff is in open literature.
Why has nobody done this? Because… they haven’t.
That’s the answer. There's no magic stopping anyone, and it's disturbingly easy. Very rough difficulty estimate: making a novel disease (or weaponizing a current one) that could kill millions is about on par with clandestinely cooking LSD. So harder than cooking meth, but we know labs have cooked acid so it's very possible. It gets easier if you're an unhinged fanatic and don't care if you kill yourself with your own plague, which means you can skip all that bunny suit nonsense.
AI boosts bioterror risk a little, I suppose, inasmuch as it might help you learn. It doesn’t affect atomic risk much since the bottleneck there is materials. Once you have enriched weapons grade material a gun type weapon can be made in a machine shop with designs available at a college library.
As for ASI becoming sentient and exterminating us, I can’t give it zero probability because there's unknown-unknowns. But I’d rank it far lower than the risk of extreme climate tipping points (e.g. the clathrate gun), old fashioned bioterror, or atomic war. And like I said, advanced AI could be used to help us with climate change if it can help us crack better energy tech.
How does that prevent AI from becoming superhuman in all intellectual domains, thus creating ASI? We set "become good as chess" as the metric and it became superhuman in that domain.
> The problem is what happens when you max that out.
I'm not sure how you conceptualize "maxing out" the intelligence metric. Again, taking chess ability as a proxy metric for intelligence, there is no reason to believe we have maxed out chess performance, but AI is already far superior to humans. And better bots are created all the time. There is no need to design a new metric to improve chess performance. The old "How many currently existing players can I beat?" is good enough. Also, why couldn't it design better metrics after achieving superhuman intelligence?
But even if we suppose there is some kind of fixed point limit to this process, it would still be far above human level. That is all that is required for ASI.
Regarding X-risk: the point is that it becomes easy even for people who, unlike you, haven't studied biology. For things such as atomic war, the AI does not necessarily need to acquire the materials. It can access them digitally by hacking the weapons systems, possibly in collaboration with some human actors. Or maybe it spoofs detection systems causing countries to fire upon each other. Generally, it seems to me that the barrier to entry for bad actors to cause these scenarios is decreasing. Whether these scenarios are more likely than D-risk I don't know.