But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.
I might need to check out DeepSeek more. I had no idea the difference was this obscene. Makes me wonder if something's off with the benchmark. A 70x cost reduction vs. Fable seems too good to be true.
GPT 5.5 Pro was ~230x at almost $23 per task.
DeepSeek is my go-to when I need an API, and local Gemma 4 won't do because it's either too slow or not capable enough. DeepSeek isn't at the frontier but it's good enough for a lot of things, very cheap, and quite fast. Flash is even faster and cheaper, and still better than anything I can host locally.
Unless the cost per token is prohobitedly high, people can often try the model out themselves and make a subjective judgement of how effective and efficient is it at solving tasks they usually deal with, using their setup.
In my own experience, Fable is more token efficient than opus 4.8 with a higher likelihood of completing tasks correctly or at least with minimal corrective work. Opus regularly struggled to gather the correct context and reason effectively about what it had gathered.
GPT-5.6-sol crushes fable in speed and token efficiency and is clearly superior across many tasks that matter for me.
I also find all models from anthropic after opus 4.6 to suffer from the same ai slop language that long plagued OpenAI and seems to have been reduced drastically in 5.6