Agreed! And with all the gaming of the evals going on, I think we're going to be stuck with anecdotal for some time to come.
I do feel (anecdotally) that models are getting better on every major release, but the gains certainly don't seem evenly distributed.
I am hopeful the coming waves of vertical integration/guardrails/grounding applications will move us away from having to hop between models every few weeks.