None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues
The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems
It literally seems like they are doing just that, and the agents are just finding holes in that.
Anything beyond baseline would be observable- silence, malformed packets, too much egress, unusually large packets, etc
If you are using a firewall, you are already doing it wrong. Don’t list thinks to block - list things to allow and make that list small.