Opus 4 beat all other models. It's good.
Opus 4 beat all other models. It's good.
If a model is really that much smarter, shouldn't it lead to better first-attempt performance? It still "thinks" beforehand, right?
There's some interesting takeaways we learned here after the first round: https://www.tinybird.co/blog-posts/we-graded-19-llms-on-sql-...
Even if you don't care about racial politics, or even good-vs-evil or legal-vs-criminal, the fact that that entire LLM got (obviously, and ineptly) tuned to the whim of one rich individual — even if he wasn't as creepy as he is — should be a deal-breaker, shouldn't it?
I wonder how much the results would change with a more agentic flow (e.g. allow it to see an error or select * from the_table first).
sonnet seems particularly good at in-session learning (e.g. correcting it's own mistakes based on a linter).
Is there anything to read into needing twice the "Avg Attempts", or is this column relatively uninteresting in the overall context of the bench?