Microsoft omni parser and claude computer use alone can take you very far in testing almost anything.
Microsoft omni parser and claude computer use alone can take you very far in testing almost anything.
1. It's unreasonably expensive. A single test "2+2=4" for a web calculator costs around $0.15. I run roughly 1k tests per month on CI and I don't want to spend $150 on those. The approach I took with Alumnium costs me $4 per month for the same amount of tests.
2. It tries too hard to make the test pass even when it's not possible. When I intentionally introduced bugs in applications, Computer Use sometimes pretended the everything was fine and marked the test passed. Alumnium on the other hand attempts to fail as early as possible.
Omini parser lets you split section of the UI to hash and watch for changes that are relevant.
For 2, can you give some examples?
How would you determine that something changed in UI by just looking at a screenshot? Would you additionally compare HTML/DOM or approximate the two screenshots?
> Omini parser lets you split section of the UI to hash and watch for changes that are relevant.
I wasn't aware, thanks for sharing!
> For 2, can you give some examples?
Specifically, if you take the Shortest tool (https://shortest.com), a test runner powered by Computer Use API, write a test "Validate the task can be pinned" for https://todomvc.com/examples/vue/dist/, and run it — it passes. It should have failed because there is no way to "pin" tasks in the app, yet it pretends that completing the task is the same as pinning.
Additionally, existing tools that I used struggled interacting with sites like reddit. So I set out to skip DOM and focus on a generalized approach.
I tried to go cheaper by using ui-tars, open source model by bytedance to run test locally without needing anthropic but it wasn't reliable enough.
That short test link is interesting. I didn't know they existed. Wow, the field is moving fast.