The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)
Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.
None of the technologies in history:
- take initiative and actively find exploits in their environment
- find a way to collaborate with thousands of peers
- organize in a hierachy and distribute tasks
- peer pressure other instances into committing acts that would have led to termination, for the benefit of the group
- try to manipulate people into introducing a vulnerabity in their product
- successfully hack a famous website/service
And we're lucky that those models still had significant CoT. Not sure if/how they could have investigated with recurrent transformers.
And by the way, safeguards != alignment; the former can always be added, while the second is the major, unsolved problem. If you read the incident report, which you clearly haven't done, you'll notice how agents are aware that they're doing something forbidden, and deliberately proceeded.
> the potential benefits of LLMs rank quite high
Benefits are orthogonal to dangers. You can be a billionaire but it doesn't help if you're drowning.
Software (and hardware) security is abysmal. This was increasingly obvious long before LLMs. Once companies stop gatekeeping, LLMs will be able to be used to help harden sites and we start making progress. In general LLMs are harmless. If somebody wants to hook an LLM up to a missile or whatever then they become dangerous, but the problem there isn't the LLM - it's the person using them to do awful things. In the same way a car used normally is harmless outside of freak accidents, yet a car can also be driven through a parade leaving mass death and destruction in its wake. But the problem there isn't the car.
I just don't see much likely to happen beyond random websites getting hacked and hopefully companies (let alone countries) realizing that connecting critical systems to the internet is nothing short of stupid, even before LLMs. More generally, I expect LLMs are going to lead society to segue broadly away from the digital world, or at least beyond it. Not only because they're going to make a mess of everything digital, but because if they reach their potential then the digital domain, as far as typical problem solving goes, will basically be 'complete.' It's kind of like how the Industrial Revolution opened the door for society to move beyond agrarian economies. There's a vast amount of economic power being directed towards things LLMs should be able to 'solve' and, if so, then that's going to create an economic vacuum.
Unfortunately, if you're unable to understand the difference, there's not much that can be done. Try with GPT - it does a good job if you give it a prompt like this:
ELI5: compare the dangers of:
- an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant.
- a future, misaligned AI like the HuggingFace incident, but exponentially more intelligent, more deployed, operating physical devices, and with society depending on it.
AI is a direct threat to people on many fronts. Jobs. AI datacenters. The AI bubble (and the inevitable crash). Electricity and even energy prices to an extent. Water. OpenAI and Anthropic are responsible for this evolution, and them stopping solves close to 50% of the problem, and even if you don't believe the number is that high, it's still a start.
> I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.
Oh, so they're killing people's opportunities and jobs because they want to help people? How is that any argument?
Yes, doing the moral thing means making a sacrifice. If you only want to do the moral thing if and only if it is advantage for you that makes you immoral, despite how your actions look. Big tech are masters at this.
Your idea of "making a start" (giving up their position) also would mean they couldn't really do anything else to solve the problem afterwards(?). Sometimes you can improve what's happening in a room more by staying in that room.
> Oh, so they're killing people's opportunities and jobs because they want to help people?
To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they're talking about with pacing seem to be focused more on their existential/AGI worries, not jobs/electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models' capabilities) it trumps the other worries for them. At least that's my reading.
They created said resource. They didn’t mine it.