Have you used 10b/5b tokens over the course of a week or over the course of a month?
The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.
Anthropic caching must be better because the cache rates are better on Claude models.