If you "don't touch the stuff", you are probably relying on your initial impressions from 2022 to evaluate models in 2025.
Also, predicting the next token with high accuracy may demand very high degrees of knowledge and reasoning. As an extreme example, please predict the next 1,000 tokens in this sequence: "ABSTRACT. In this paper, we show a mathematically elegant unification of quantum mechanics and gravity, which makes testable predictions. Testing these predictions shows that the quantum gravity model provides previously unexpected results accurate to 1 part in..."
To predict the rest of that paper, it helps to actually come up with a workable model of quantum gravity. Which no current LLM can do, happily.
But lots of current models are good enough at "predicting the next token" to solve high school honors math problems that they've never seen before. They can apply the chain rule, factor polynomials, double-check their work, backtrack, etc. Similarly, current-generation coding models are perfectly capable of reading compiler error messages, and generating diffs that fix the underlying problem. Current-generation summarization models are capable of reading several scientific papers, extracting the key concepts, and turning them into a fairly serviceable podcast.
All of this happens because a (1) in order to predict the next token better, LLMs actually build thousands of specialized models that do things like "keep track of the state of a chess board" or "recognize lions in photos", and (2) transformer models are sufficiently powerful to model many problems well. So a big current-generation LLM is basically an ensemble of thousands of domain models glued together with a language model and a bunch of feed-forward layers.
Now, none of the outputs from LLMs are of super high quality, compared to experienced humans. And there are deep reasons for that. But if you happen to have problem where medium-competence AI "slop" is actually beneficial, then yes, LLMs can actually be of real-world use.