We switched to a 5x cheaper LLM. Our costs went up
gitar.ai
gitar.ai
Three things bit us: finish_reason semantics differ between "compatible" providers, the model retried identical failing tool calls instead of adapting, and provider failover invalidated prompt caches on both sides.
Curious if others have hit similar issues.