For whatever reason, the AI companies are (or at least were) not doing this kind of classification online during their testing runs, and instead just checking transcripts after the fact. This is more clear in the Anthropic reports about their incidents, for example:
"The earliest incidents date to April ... We began our transcript review on Thursday, July 23, and stopped all cyber evaluations the same day after identifying transcripts where Claude may have accessed the internet"[1]
I agree that it is crazy and negligent! I don't think it's good for their business, though -- who wants to use a model that will just cheat instead of doing the job you asked for?
> My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.
At some point, if you are making an LLM in order to use it for useful work, it really benefits you to give it broad network egress.
[1] https://www.anthropic.com/news/investigating-incidents-cyber...