Yes, but the generalists are not routinely improving across all domains. The large labs are really focusing on agentic use, so I imagine that creative writing has deteriorated considering how distinctive Claude's writing style has become. Or I recently had an image-parsing task, and I was excited to try Qwen because I heard it had gotten a lot better at agentic tasks, but it failed my image-parsing benchmark.