I have had AI models perform well on obscure use cases. Even two years ago, let alone the latest models, I've had ChatGPT generate correct code for bespoke requirements that I could not find anywhere else on the Internet. I actually spent a lot of time trawling GitHub and even SourceForge to convince myself this wasn't just a case of "stochastic parrots" regurgitating existing code from some forgotten corner of the Internet. (This also served as an exercise in researching prior art for what I as trying to do.)
In one case it made sense of my own gnarly, hacky, organically-grown, prototyping code that I myself had lost all mental context on and devised the correct solution to my prompt. Two minutes of prompting saved me an hour's worth of software archaeology.
Each time, the model seemed to "understand" the requirements, "grok" the existing code, identify the relevant concepts needed, and apply them to the task at hand. I certainly did the hard work of researching the problem and figuring out the high-level approaches, but the AI was very good at generating fairly complex code to implement those ideas.
While there certainly is a ton of overhyping going on, especially from the "AGI-is-nigh" quarters, the hype is compelling because these models really do succeed at useful tasks much more often than not.