I, one of those end users, absolutely use these benchmarks as heavy input considerations. Indeed, the vast majority of people do. "Completely meaningless" is just nonsense, of course, and while it doesn't perfectly map to every use, there is a pretty good correlation with suitability for specific tasks.
I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.
Not to mention that the linked page includes a pretty broad list of specialization benchmarks.