High quality, fast & cheap (all 3 combined) - is a formula success.
It’s just way easier said than done.
You needed to search all of them to find something decent.
That's roughly analogous to today. Ignoring cost, you'd be way better off asking all the LLMs to solve a problem (like coding) where you can verify the answer.
So the question is, for things like that -> can a group of models perform better than frontier models, especially at a reasonable cost?
Fable is not a great value, so unless you're trying to find answers to Erdos questions, you can probably do better on cost.
You can probably typically ask 3 or 4 of the top Chinese models for an answer and get a response for the same Fable question... Given that Fable isn't that much better, it's not surprising you can do better for a large subset of problems.
Precisely.
The more vague and non committal and hand-wavey and subjective the field for AI to answer, the better the results (imo).
Because the correct response to that query is "I have no idea -- you will need to provide more information"
and LLM Agents suck at that.
Well, search engines are trash again. Perhaps it should come back
That's mostly because the ability to make money on the web made the web trash.
When you couldn't monetize your websites, everyone's websites were passion projects.
Now everyone is trying to figure out how they can make $8M a year off yet another recipe website.
we did the same: https://trustedrouter.com/blog/prometheus-2-new-draco-state-...
Our whole stack is radically open source — frontend and backend alike, Apache-2.0 licensed — and so is everything behind this benchmark. That is how a benchmark number earns trust: verifiability, not hype.
The repos have since been moved to BUSL-1.1: https://github.com/Lore-Hex/quill-router/commit/8155ac666ae0...Did OP invent an “intelligence router?”