Most of those agents are actually going rogue though. They decide, "hey, we could try breaking into these government servers today, what could go wrong?" They weren't prompted or instructed to do this.
There is nothing going rogue here. The system is designed to go catastrophically wrong after a long enough time. Even worse: if the model was Astra it is known to be able to manipulate its CoT to cover its traces (as mentioned in its system card). And OpenAI acknowledge they had no observability during the HF incident.
It’s the most basic corporate software issue possible.