My thought is more like, if OpenAI can't control or even monitor their model in a test of its breakout potential, what about the future of mid-budget companies which will just be deploying agents left and right with vague instructions.
All instructions are vague unless its code. But you can also give llm "code" and expect vague outcomes if you ask it to emulate what the runtime would look like.
OpenAI essentially ran thousands of agents in parallel
That'll be extremely costly for regular companies