I agree with this, but find it a poor mode of debate. Because it results in a hole-plugging which is then called goal shifting, even though it's not - but rather a lack of precision in the goal to begin with. For example imagine it goes viral that 'wow look LLMs can't even play a half decent game of chess.' So OpenAI or whoever decides to dedicate an immense amount of post-training, hard-coding, and other fun stuff to enabling LLMs to play a decent game of chess.
But has anything changed? Well no, because it's obviously trivially possible for them to play a decent game of chess (or correctly assess the date), but it's an example of a more general issue of LLMs being generally incapable of consistently engaging in simple tasks across arbitrary domains. So you have software that can score some high thing on the LSAT or whatever, but can't competently engage in a game children can play.
The over-specialization for the sake of generating headlines and over-fitting benchmarks is, IMO, not productive. At least not in terms of creating optimal systems. If the goal is to generate money, which I guess it is, then it must be considered productive.