You raised a good point, what's a good metric for LLM performance? There's surely all the benchmarks out there, but aren't they one and done? Usually at release? What keeps checking the performance of those models. At this point it's just by feel. People say models have been dumbed down, and that's it.
I think the actual future is open source models. Problem is, they don't have the huge marketing budget Anthropic or OpenAI does.
It’s not always for the best mind you, but I do think major changes to the GUI or UX do deserve a version indicator given they are relatively infrequent. Without some requirement of course.
Nobody buys Oracle because of the UI or upgrades because of a new UI. It’s the underlying tool that matters and if my ugly version is working there is no incentive to upgrade to a prettier version.
It doesn't matter if a model is e.g. 30% cheaper to use than another (token-wise) but I need to burn 2x more tokens to get the same acceptable result.