I maintain the notion that LLMs are just what all NFT grifters moved to after the NFT fad died.
I don't care whether the LLM can solve this or that mathematical previously thought unsolvable theorem. What I do care about is can it write good, maintainable code. And every new release of frontier models - they get better at it.
Good output depends on good input (prompt), and a good set of available tools the model can use to verify their work. If given this, nowadays really you need to try to get the model to output garbage.