https://dev.meta.ai/docs/getting-started/pricing-rate-limits
https://dev.meta.ai/docs/getting-started/pricing-rate-limits
Grok 4.5 has a relatively high $0.50 per 1M cached input token rate, compared to $0.15 on this model.
Grok 4.5 cached input costs the same as Opus 4.8 cached input, which is going to make it a lot more expensive to use for multi-turn coding than many would assume from the $2/$6 headline numbers they led with.
There's a further sting in the tail, Grok 4.5 is only $2/$6 for the first 200k of context. Go above that, and the pricing is $6 / $12 - and you're still capped at only 500k context anyway.
Here's the xAI pricing on OpenRouter:
https://openrouter.ai/x-ai/grok-4.5?endpoint=0e927811-b1a8-4...
Compare with Grok 4.5 which came out at $2/$6 but then quietly charges $0.50 per 1M cached input tokens. That's as high as Opus 4.8!
If they have a really good model, it makes sense to subsidise it, to gain users, before they align prices with competitors.
I really dont see how anyone's willing spend more than $1.50 per mm output. Let alone $15-50. Does anyone actually pay for usage based billing as a consumer?
https://platform.claude.com/docs/en/about-claude/pricing
Model Base Input Tokens 5m Cache Writes 1h Cache Writes Cache Hits & Refreshes Output Tokens
Claude Fable 5 $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok
Claude Opus 4.8 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok
Note Fable costs $50 MTok and Opus 4.8 costs $25 / MTok.
Even with usage based billing I'm below $1 writing code all day.
same, pi as a harness works best, better than claude code and open code.