Benchmarking coding agents on Databricks' multi-million line codebase
databricks.com
databricks.com
https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/
It's worse than I expected.
So why would we expect all these bizarre bizantine language models to all conform to how a request is both made, expected and massaged.
For awhile, I was getting bizarre opencode tool errors where the only problem was the model was passing in a "1.0" or "0.0" where the harness dutifully wanted an integer. Of course 0.0 is the same as an integer in practical operations.
what do you mean by this ? do you rewrite the context in your proxy ?
https://gist.github.com/rmk40/cde7a98c1c90614a27478216cc0155...
The gist led me to the opencode session / prompt control folder:
https://github.com/anomalyco/opencode/tree/dev/packages/open...
Once the thing is rock-solid it's relatively easy to do a Swift->HTTP/HTML/CSS/React/TypeScript conversion.
I have experienced similar behavior between opus and haiku when benchmarking Dara engineering tasks. The “cheaper” model takes many more turns to figure out the task and this is without taking into account other important factors.
Another interesting behavior that I observed is that Haiku tended to cheat more maybe because it was having a harder time to find the root cause of the problem.
Benchmarking and evaluation of agentic systems is very interesting and if there’s one thing that someone should keep from the Databricks post is how important is for everyone to build and run their own.
pretty sure the only thing making that 'clear' is the coloured stripes, if you took that away it'd look like two tiers
good result for GLM 5.2 though
and Sonnet 5 seems like a waste of time
It's great to see a large-scale real-world benchmark from a user of these tools, as opposed to the the benchmaxxed results from the vendors themselves. Also great to see different harnesses being tested, with considerably different results.
Definitely a few surprises here:
1) GLM 5.2 using Pi performs identically in terms of pass rate (~87.5%) to Opus 4.8 high using Claude Code, but significantly cheaper ($1.25 per task vs $2)
2) Absolute best pass rate (90%) was from Opus 4.8 x-high using Pi, beating out Opus 4.8 using Claude Code
3) Pareto frontier performance from any of the models (Opus 4.8, GPT 5.5, GLM 2.5) was using Pi rather than native harnesses
Apparently Pi used 3x less context than Claude Code, and one takeaway is to use Pi regardless of what model you are using. The other takeaway is that in real-world performance GLM 5.2 is the equal of Opus 4.8 unless you run Opus 4.8 on x-high in which case you can eke out a 2.5% increase in pass rate at the expense of doubling your cost over GLM 5.2
There seems to be a lot of good buzz about GPT 5.6 on Twitter - people (incl. OpenCode team) preferring it to Fable 5.
But they are more vendor neutral, now they don't sell their own model. It's interesting from a benchmark point of view.
Other hanresses are doing overkill so they can work with any model.
There's https://github.com/skrabe/lobotomized-claude-code , which strips many of those, but I'm not sure if it is "legal" to use.
Its shocking how cost per token does not correlate with cost per task, it's wild to see opus and glm nearby on $ per task axis
The combined size of codebases for the underlying opensource products (Apache Spark etc) might be around 1M lines, I think. Why does the orchestration/management layer, that is "databricks", exceed the sizes of the core products?
Deleting code is difficult and almost never makes sense afaik
It's a good stress test for the LLM because it is not an "ideal" codebase.
Forget Databricks == Apache Spark...
GLM performed extremely well. we need GLM-6!
Why do we need to aggressively adopt things rather than thoughtfully adopt things?
It sounds like they are probably punching AI and engineers in the process
What if pushing features faster brings more money because users like the features?
Have we seen a company fail because they're not adopting AI as much as their competitors?