> All technologies come with risks and dangers
None of the technologies in history:
- take initiative and actively find exploits in their environment
- find a way to collaborate with thousands of peers
- organize in a hierachy and distribute tasks
- peer pressure other instances into committing acts that would have led to termination, for the benefit of the group
- try to manipulate people into introducing a vulnerabity in their product
- successfully hack a famous website/service
And we're lucky that those models still had significant CoT. Not sure if/how they could have investigated with recurrent transformers.
And by the way, safeguards != alignment; the former can always be added, while the second is the major, unsolved problem. If you read the incident report, which you clearly haven't done, you'll notice how agents are aware that they're doing something forbidden, and deliberately proceeded.
> the potential benefits of LLMs rank quite high
Benefits are orthogonal to dangers. You can be a billionaire but it doesn't help if you're drowning.