HNHacker News
TopNewBestAskShowJobs

Edward40

3 karma · joined September 2, 2026

submissionscomments
Edward40··on Errand – open-source Grok Bot and Muse alternative, built in a week
I am really impressed by Runta's infra. Integrating remote desktop was basically effortless.
Edward40··on Show HN: Open-source AI teammates with their own computer
Wild that we built an open-source Grok Bot alternative in just one week on top of Runta. Remote desktop integration was practically plug-and-play.
Edward40··on Dive into FrontierHarness Eval: Claude Code cost 5.6× for the same pass rate
Claude Code is an opinionated harness. I’d love to see whether neutral harnesses can challenge the monolithic harness narrative.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
We'd rather not mess with system prompts, we just eval them as shipped to keep the results reproducible.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
I also prefer using vanilla Pi over Oh My Pi.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
Thanks! V1.1 is set to add more harnesses and evaluate a broader range of models. We're moving to a full harness × model matrix to uncover interaction effects and release a complete harness-model compatibility map.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
Good points! That’s a known limitation of our v1.0 benchmark with Claude Code. For v1.1, we're expanding both harnesses and models to evaluate the full harness × model matrix. This will highlight interaction effects (showing which harnesses suit which models best) and let us publish a full compatibility matrix. We shared more details on our roadmap in our blog post (https://runta.com/blog/introducing-frontierharness-eval/).
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
We initially tried using the mean (average), but a few extreme outliers caused Claude Code's cost to look far higher than it typically is.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
Version 1.1 will add more harnesses and evaluate a broader set of models. Rather than holding the model fixed, we plan to evaluate the full harness × model matrix to reveal interaction effects (which harnesses work best with which models) and produce a harness-model compatibility matrix.
Edward40··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
The Exo harness is the most interesting one in the benchmark, as it has the potential to complete tasks with fewer turns and cost.