That's what they signed up for when established a hard dependency on an subscription online-only LLM model.
And self-hosting would probably have been more expensive unless they had massive volume: Deepseek V3.x would have been the comparable open weight model for the performance and isn't that cost effective until hosted across multiple nodes with large batch sizes
If you want to be sure to be in control, then host it yourself
Because fuck you, that's way. /s
There's a massive amount of contempt from a lot of software businesses for their customers, and it's just getting worse with LLM models.
It's not because a model performs better in some applications (often by fine-tuning to get better scores at specific tests) that it is better across the board or that we have to believe the company releasing the model with a high number 3 > 2 so that it is commonly accepted as better.
Pushing the reasonnning further: f you need an Opus level performance then not accepting GPT 3 isn't a smell.
What's often missing is the actual engineering. If you have evals for your use case, you use these to adjust to the new model. DSPy has great prompt engineering constructs it ships with.
The smell is all about vibing, where a model feels better because the structure of the answers is more familiar to a person, instead of engineering where given constraints and input/outputs one is using/bending the system to fulfil the requirements.
However, they are aggressively deprecating them (OpenAI is as well), and replacing with newer models. These newer models are all reasoning models, and importantly, only bear the flash name. They are not fast. And they are very expensive!
But this could be framed as 'getting attached to an API revision when a new one is available'...
I can see it both ways, tbh.
I tried 3 flash for months and it didn’t work using Googles own vertexai integration because it’s been in preview mode for months.
Not wanting to pay significantly more and do a bunch of rework isn’t a smell.
They left a large gap in their new pricing vs the prior generation, and if you had a working use case that sucks. The model is >99% reliable for my use case so there’s nothing to gain from a smarter model.