> Block any and all network requests.
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.