87% vs 56% on Webvoyager
58.1% vs 36.2% on WebArena
38.1% vs 22% on OsWorld
These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)
87% vs 56% on Webvoyager
58.1% vs 36.2% on WebArena
38.1% vs 22% on OsWorld
These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)
The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with some browse-use agents that are SOTA and they can still get very easily confused with more complex tasks or unfamiliar interfaces.
You can see some of the examples in OpenAI's blog post. They need to quite carefully write the prompts in some instances to get the thing to work. The truth is that needing to iterate to get the prompt just right really negates a lot of the value of delegating a one-off task to an agent.
No. It's not matching them, it's clearly exceeding them. The previous post provided the numbers.
In WebArena, Operator does 58.1%. Previous SOTA for browser-use agents is 57.1%. In WebVoyager, Operator does 87.0%. Previous SOTA for browser-use agents is the exact same.
See here for details: https://openai.com/index/computer-using-agent/