I can speak for my area of expertise - multilingual capabilities. Some SOTA models are making huge strides in their support of various languages, and increasingly they understand and can produce text in languages where GPT-4 era models were absolutely lost. These are probably from a combination of richer training dataset and architectural improvements (more parameters?).
I posted about this here if you're interested: https://news.ycombinator.com/item?id=47847282
Now that doesn't necessarily mean that models are also getting substantially better at English or other major languages. They likely are to some degree, but we've reached a point with major languages where core linguistic proficiencies are covered, and what's left is the more squishy part: style, tone of voice, ability to use different registers naturally, or what some people would call linguistic taste. But that's much harder to measure and therefore trickier to provide evidence for.
Hope this helps.
Edit: typo, clarification