Hmm, am no LLM expert, but agree with you that the models themselves for the individual subject domains seem like they're starting to reach their peaks (Writing, solving math, coding, music gen...) and the improvements are becoming a lot less dramatic than couple of years ago.
But, feel like combining LLM's with other AI techniques seems like it could do so much more...
... As mentioned, am no expert, but seems like one of the next major focuses on LLM's is on verification of its answers, and adding to this, giving LLM's a sense for when its result are right or wrong. Yeah, feel like the ability for an LLM to introspect itself so it can gain an understanding of how it got its answer might be of help if knowing if its answer is right (think Anthropic has been working on this for awhile now), as well as scoring the reliability of the information sources.
And, they could also mix in a formal verification step, using some form of proof to prove that its results are right (for those answers that lend themselves to formal verification).
Am sure all this is all currently being tried. So any AI experts out there, feel free to correct me. Thanks!