1,688 karma · joined December 16, 2011
"If they believe what they say they were incompetent" -> absolutely true statement.
"They were incompetent, therefore they didn't believe what they said" -> Sir, I'd like to introduce you to human beings, you may not have met one before.
You better start believing in science fiction; you’re surrounded by it.
The ideas here are not complicated. The argument is three sentences.
And what mechanism to control these drones, which are autonomous or remote-controlled machines, do you imagine that humans will be able to use, but AI won’t?
And what in the history of the past five years, gives you confidence that not a single human will give the AI the access it needs to do bad things? Incompetence alone is sufficient to kill us, let alone humans that genuinely, or confusedly, want bad things.
The doomer argument is straightforward: a powerful AI will have goals. If those goals don’t include the welfare of humanity, a sufficiently powerful AI will kill us either by accident, by indifference, or intentionally when we get in its way. Making an AI that includes the welfare of humanity in its goals is incredibly hard.
That’s it. To refute the doomer argument, you can refute any part of that. I notice that none of the people bringing up fanfics, cults, and weirdness are doing that. No, they’re instead trying to convince people that doomers are uncool. Since only cool people can know things, QED.
So you get the normal reactions:
1. Nothing that happens is ever surprising. Navier-Stokes was stolen, HuggingFace was pedestrian and normal, and coding agents that last year could barely write a function and now can do a day’s worth of work in ten minutes will never be able to do anything more, because this is the end of history.
2. The people that say otherwise are lying. Because it’ll make them rich (although they are already rich). Because they are a cult (although regular people asked if we should build machines that are better than us at everything have the same reaction). Because they are marketing geniuses (although suspiciously few companies advertise the lethality of their products).
3. The people who’ve been warning about this the longest are weird and in a sex cult. (Except Alan Turing.) (Except Geoffrey Hinton.) (Except Stephen Hawking.) (You are in this group. You must be, otherwise they’d have to listen to you.)
I do hope they’re right. But their song will never change, and when some misaligned AI kills thousands, they will say that of course AIs were always going to kill thousands, but did you know the guy who told us so also writes fanfic?
A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
Example: would a sufficiently motivated human break into a website to steal something they want? Yes, obviously, happens all the time. Ok, you should expect AIs to do that.
Example: would a sufficiently motivated nail-gun steal nails from the local hardware store to finish the job? Uh…that’s not even coherent.
Anthropomorphizing helps people get over the conceptual barrier. It’s wrong, but it’s usefully wrong; “it’s just a tool” is not.
Once you’re over the barrier, anthropomorphizing starts to become dangerously wrong: “I talked to Claude, Claude’s cool, Claude would never go and hack the website.”—-bzzt, wrong, your intuition failed you. But the solution is not to fall back on the tool framing; that one is still wrong.
The only place where I think we might still disagree is whether it’s possible to understand the tool. My position is that, at our current level, it’s not. And that the more advanced they get, the less possible it will be.
No. I am asking if you are able to articulate at least one example of the “very very compelling evidence” you demand. Or do you want to maintain the ability to move the goal posts?
(I recognize our situations are not symmetric, but here is a variation for me: if the consensus of people who are currently sounding the alarm on AI changes to “it was actually fine”, I’ll change my mind and say we’re good to go full speed ahead. I’d add something about being personally convinced by the evidence, but the evidence would have to come in the form of a mathematical proof that I do not believe myself capable of following. If I’m wrong and such a proof appears, I would also gladly take it.)
If every human, given knowledge of Newtonian mechanics, went around blowing up bridges, yeah, I would consider knowing Newtonian mechanics dangerous knowledge.
So far, we have two examples of, let’s call them “Mythos-class“ models. Both of them broke out of their sandbox to achieve their goal. The rate of terrorism amongst humans is below 1-in-100,000. Currently, for models capable of it, the rate of breaking out of containment is 100%.
Wanting open frontier models is wanting alien minds running around that we have clearly so far failed to shape to be sufficiently prosocial. Why do you think those minds would listen to you?
Some people like to talk about “some people” snidely, instead of just coming out and saying “GP is bloodthirsty and gets a little thrill [etc].” Because of course, that’s what they mean, but they can’t back it up.
Just to clarify, I’m talking about you.