"Better" is not one dimensional across all use cases even if model capabilities are improving in aggregate.
e.g. If someone said "this is the worst they'll ever be" in response to some writing with obvious LLM cliches in 2024, I'm not convinced that prediction was actually correct.
The focus of OpenAI/Anthropic pivoted aggressively to the agentic performance arms race instead of making a more human sounding chatbot so regressions in writing ability aren't really a concern anymore if agentic benchmarks improve.
The first time I heard a recommendation to use Claude was specifically because it sounded much more "human" and natural than ChatGPT. Fast forward to now and idiosyncratic Claude-isms repeated every other sentence and its convoluted verbosity has become a widely mocked meme.