The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
Why?
W what?
There isn't much compelling reason to use these unless you are just averse to giving money to openai/altman. A $20 codex sub gives you ~$150 of luna use per weekly limit, while there isn't any good subsidized options for chinese models at all (and the few who were subsidizing, like opencode, rugpulled by reducing monthly limit to $60 to $15 with no notice to users).
The answer is just a second $20 subscription.