What or other opinions on how representative the AA rankings are of real-world performance? Any better indicators?
What or other opinions on how representative the AA rankings are of real-world performance? Any better indicators?
[0] https://artificialanalysis.ai/models/qwen3-8-27b?models=deep...
They have simply decided to not train the model in some areas such as world physics
I have consistently noticed Opus 4.8 and GPT-5.6 far outshine the Chinese models. Gemini is sort of middle of the road, Grok is better than Gemini but not really close to Opus/GPT. OAI & Anthropic still remain unbeaten by a wide margin in my eyes.
Surely most of your use-cases are not novel tasks that combine obscure domains.
It seems to me the real way to evaluate the value of a model is how it performs in your real-life workflows.
Maybe my Greek idea sounded too high falutin' or simply seemingly clever (I give an example below -- try it out!).
So here's something way simpler that Qwen3.8 27B does not get; only GLM-5.2 and K3 do.
**
Analyze the 2 structural (not semantic) patterns in this text:
1. Morning revient.
2. Birds saluent Morgenlicht.
3. We suivons Waldwege toward maison.
4. Rain tombe plötzlich; we cherchons Schutz beneath sapins.
5. Night vient langsam; we trouvons Wärme near le Feuer, sharing quelques Geschichten together.
**
The answer should get not just the obvious cyclic E-F-G pattern but also the word counts being Fibonacci. Surprisingly few models get this. The only way I got Qwen to do this was on the 2.4T model, with extensive prompt scaffolding. Claude (Opus & Fable), Sol, Grok etc. got it on the first attempt. (All models, all attempts max reasoning level.)
Can you share an example of the Greek-history puzzle prompts you're talking about?
**
The stone remembers not what the cities gave, but what the goddess kept.
Begin with the first reckoning, under Ariston.
Find those who carried the Greeks in their name and the silver in their care. Take their name as we give it to them now, in capitals. Thirteen marks.
The goddess kept one from sixty. She asks the same of every mark: give each its ordinary alphabetic number, divide by sixty, and keep what she could not take.
Each remainder walks with the next. The last returns to the first.
The first of each pair turns upon itself; the second joins it; then comes the place where the pair began.
Eleven takes its fill. Keep what remains.
Raise eleven courses beneath the thirteen marks.
In course (r), beneath mark (i), add the course to what Eleven returned. Reduce again by Eleven. Cut the stone if the result is the remainder belonging to mark (i), or to the mark walking beside it.
Count courses from zero.
The mason turns where the stone turns: the first course goes with the writing, the next against it, and so on.
A cut is `#`. Stone is `.`.
Restore the fragment.
**
Answer is:
........#....
.#.##.......#
...#.........
........#....
#.........#..
#.....#.#....
.#.#.#......#
..#..#.#..#..
.........#...
.#...........
.#....##.....
I do like the output from qwen when I get it, but honestly I haven't been impressed enough with it to put up with the downsides.