For example, the latest Gemini 2.5 Flash is known as "google/gemini-2.5-flash-preview-09-2025" [1].
[1]: https://openrouter.ai/google/gemini-2.5-flash-preview-09-202...
This is also the case with OpenAI and their models. Pretty standard I guess.
They don't change the versioning, because I guess they don't consider it to be "a new model trained from scratch".
In all seriousness though, their version system is awful.
That "example" is the name used in the article under discussion. There's no need to link to openrouter.ai to find the name.
That's why Google names it like this, but I agree its dumb. Semver would be easier.
semantic versioning works for most scenarios.
This is all solved for a long time now , llm vendors seems to have unlearnt versioning principles.
This is fairly typical - marketing and business wants different things to do with version number than what version number systems are good at .
This is the entire premise behind the cloud, the reason it was Amazon did it first, they had the largest workloads at the time before Web 2.0 and SaaS was a thing.
Only businesses with large first party apps succeeded in the cloud provider space, companies like HP, IBM all failed and their time to failure strongly correlated to their amount of first party apps they operated. i.e. These apps anyway needed to keep a lot of idle capacity for peak demand capacity they could now monetize and co-mingle in the cloud.
LLMs as a service is not any different from S3 launched 20 years ago.
---
[1] It isn't, at the scale they are operating these models it shouldn't matter at all, it is not individual GPUs or machines that make a difference in load handling at all. Only few users are going to explicitly pining a specific patch version for the rest they can serve either one that is available immediately or cheaply.