I mean, what good is a prediction that is 50% accurate? If you are classifying documents for a recommendation model a "up/down" classification is barely useful, a probability calibrated classification is golden. With no calibration you have an arXiv paper, with a calibration you can build a classifier into a larger system that takes actions under uncertainty.
The generative paradigm holds progress back. You can ask ChatGPT to do anything and it will do it with 70-90% accuracy in all but the hardest cases. Screwing around with prompts can get you closer to the high end of that range, but if you want to do better than that you've got to define your problem well and go through a lot of the grindy work that you had to do with symbolic A.I. and have always had to do with machine learning. (You're going to need a large evaluation set to know how well your prompt-based solution works, and know that it didn't get broken by a software update, at the very least.)
The image that comes to my mind, almost intrusively, is Mickey Mouse from the movie Fantasia where he shows various sins, laziness most of all
https://www.youtube.com/watch?v=VErKCq1IGIU
So many of these efforts show off terrible quality control. There is a site that has posted about 250 galleries (at a rate of 2 day) of about 70 pornographic images a piece generated by A.I. At best the model generates highly detailed images including the stitching on the seams of clothes, clothing with floral prints matching cherry blossom trees in the background and sometimes crowds of people that really click thematically. Then you notice the girls with two belly buttons and if you look enough you'll see some with 7 belly buttons and realize the model doesn't really understand the difference between body parts and skin so there is a nipple that looks like part of the bra rather than showing through the bra, etc.
Then there are the hideously distorted penises that are too long, too short, disembodied, duplicated, bifurcated, pointing in the wrong direction and would otherwise be nightmare fuel for anyone with castration anxiety.
If the wizard was in charge he'd be cleaning these up, I mean looking at 150 images a day and culling the worst is less than an hour of work. But no, Mickey Mouse is in charge.
"Chat" in "ChatGPT" is a good indication of what is going on because it is brilliant at chat where it can lean on a conversation partner to provide meaning and guidance and where the ability to apologize for mistakes really seduces people, even if it doesn't change its wrong behavior. The trouble is trying to get it to perform "off the leash" at a task that matters is a matter of pushing a bubble around under a rug, that "chasing an asymptote" situation is itself seductive and one of the worst problems in technology development that entraps the most sophisticated teams, but put it together with unsophisticated people who don't think systematically and a system which already has superhuman powers of seduction (e.g. "chat" as opposed to problem solving) and you are cruising for a bruising.
I mean, a typical LLM is also logistic regression, but it's not linear.