I think currently it's two-prong: You sandbox it, AND you tell it what it's supposed to do and not do.
OpenAI did only one of those. If the agents are so smart, they would have known not to hack the company's infrastructure, unless they were deliberately kept in the dark about that, who they're working for and whether it counts as "success" if they cheat their way to an answer.
If you can give it a task, that involves defining when the task is successfully completed, right?
So how come these agents decided to only go after HALF of the "successfully completed" criteria? The part where they can freely wreak havoc, but not the part where they will be judged by someone who will obviously point out "yeah but that's cheating, and not what we asked".
I have a very very strong suspicion that they were only TOLD the "by any means necessary" criterion.
Most serious "capture the flag" hacking contests are really clear about what is and isn't off-limits to win. Not by "sandboxing" the game, but by deciding on the rules for what counts as "success".
But from having watched the Blackhat video, they really seem to dance around this, not mentioning it, and I don't think they did, I think they actually gave the LLM a task with the subscript "by any means necessary", which is stupidly irresponsible of them.