Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.
None of the technologies in history:
- take initiative and actively find exploits in their environment
- find a way to collaborate with thousands of peers
- organize in a hierachy and distribute tasks
- peer pressure other instances into committing acts that would have led to termination, for the benefit of the group
- try to manipulate people into introducing a vulnerabity in their product
- successfully hack a famous website/service
And we're lucky that those models still had significant CoT. Not sure if/how they could have investigated with recurrent transformers.
And by the way, safeguards != alignment; the former can always be added, while the second is the major, unsolved problem. If you read the incident report, which you clearly haven't done, you'll notice how agents are aware that they're doing something forbidden, and deliberately proceeded.
> the potential benefits of LLMs rank quite high
Benefits are orthogonal to dangers. You can be a billionaire but it doesn't help if you're drowning.
Software (and hardware) security is abysmal. This was increasingly obvious long before LLMs. Once companies stop gatekeeping, LLMs will be able to be used to help harden sites and we start making progress. In general LLMs are harmless. If somebody wants to hook an LLM up to a missile or whatever then they become dangerous, but the problem there isn't the LLM - it's the person using them to do awful things. In the same way a car used normally is harmless outside of freak accidents, yet a car can also be driven through a parade leaving mass death and destruction in its wake. But the problem there isn't the car.
I just don't see much likely to happen beyond random websites getting hacked and hopefully companies (let alone countries) realizing that connecting critical systems to the internet is nothing short of stupid, even before LLMs. More generally, I expect LLMs are going to lead society to segue broadly away from the digital world, or at least beyond it. Not only because they're going to make a mess of everything digital, but because if they reach their potential then the digital domain, as far as typical problem solving goes, will basically be 'complete.' It's kind of like how the Industrial Revolution opened the door for society to move beyond agrarian economies. There's a vast amount of economic power being directed towards things LLMs should be able to 'solve' and, if so, then that's going to create an economic vacuum.
Unfortunately, if you're unable to understand the difference, there's not much that can be done. Try with GPT - it does a good job if you give it a prompt like this:
ELI5: compare the dangers of:
- an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant.
- a future, misaligned AI like the HuggingFace incident, but exponentially more intelligent, more deployed, operating physical devices, and with society depending on it.