We made something called Divergence-300 @32 (and later @512) which tests actual inference across 32 tokens on a held out test (Terminal Bench, DeepSWE, Math etc)
We do plan to do larger benchmark suites though!
We do plan to do larger benchmark suites though!
The current benchmark suites that frontier AI labs use are probably a good fit, e.g.
https://z.ai/blog/glm-5.3#:~:text=Performance%20across%20com...
https://www.kimi.ai/ai-models/kimi-k3#:~:text=Performance%20...
https://www.anthropic.com/news/claude-opus-5
https://openai.com/index/gpt-5-6/
But guessing from your current benchmarks, I assume that you are severely compute-constrained. What is your time budget?