HNHacker News
TopNewBestAskShowJobs

MiguelG719

10 karma · joined September 8, 2020

submissionscomments
MiguelG719··on For Computer Use, the harness matters as much as the model
It still feels like the harness and the model need to co-evolve together
MiguelG719··on For Computer Use, the harness matters as much as the model
It changes so frequently, and the world wild web is vast so it’s dependent on your use case. In general the leaderboard reflects what we see working across a broad range of domains, but the best way to tell is to define your tasks and run the evals yourself; that’s what this is for
MiguelG719··on For Computer Use, the harness matters as much as the model
You can swap the driver/tool yourself with the evals CLI! Just use `—tool-surface <one-of-the-supported-tools>`. Or define your own using the interface
MiguelG719··on For Computer Use, the harness matters as much as the model
Token efficiency, performance, observability, and most importantly: permissions/security policies
MiguelG719··on For Computer Use, the harness matters as much as the model
Try https://stagehand.dev/dino ;)
MiguelG719··on For Computer Use, the harness matters as much as the model
Hi HN,

Over the last 2 years, we observed computer use models improving at a rapid pace and saturating benchmarks. This new benchmark replaces Online-Mind2Web with our own Browserbase Benchmark v2 that better represents the complex tasks that browser agents face in the real world. It runs against 23 models (frontier and open-weight) and 9 harnesses (Claude Code to LangChain Deep Agents) on accuracy, speed, and cost.

This new benchmark confirmed our belief that the choice of an harness is becoming as important as the choice of a model. For example: claude-opus-5 runs 74% at $1.50/task on LangChain deep agents but 71% at ~$10/task on fx.

The eval harness is a CLI you can run yourself (pick harness + tools/mcps + model, pass high-level tasks, grades with LLM verifiers, has trials/concurrency/OTEL tracing): https://github.com/browserbase/stagehand/tree/main/packages/...

Happy to get into methodology, and if you want your model or harness added, just let me know.

MiguelG719··on We made Playwright 2x faster and 80% more token efficient
We also just added Jev support https://github.com/browserbase/stagehand/pull/2952

On the act/extract/observe evals it shows promising results being extremely efficient

- act: 4.3x faster, 97% fewer LLM calls. Pass rate: 97.5% -> 98.3%. - heldout: 4.1x faster, 78% fewer LLM calls. Pass rate: 87.5% -> 97.5%. - observe: 11.1x faster, 69% fewer LLM calls. Pass rate: 75.0% -> 83.3%. - extract: 8.7x faster, 75% fewer LLM calls. Pass rate unchanged at 92%.

cost effectively 0

MiguelG719··on Fara-7B: An efficient agentic model for computer use
> but then you want to capture that in code anyway for repeated use.

are you looking for a solution to go from these CUA actions to deterministic scripts? check out https://docs.stagehand.dev/v3/best-practices/caching

MiguelG719··on Gemini 2.5 Computer Use model
are you running into this issue on gemini.browserbase.com or the google/computer-use-preview github repo?