LLMs are still prone to hallucinations, and we're only adding more data and workarounds to make it happen less often. Prompt injections limit their usefulness on untrusted inputs. They can't do precise logic and reasoning, and are too likely to follow memorized patterns instead.
We've got something amazing, way better than what we've had before, but the current architecture is still based on a fuzzy translator. It's hard to say whether this is it, and it's going to plateau at this level for a while, or whether there are more breakthroughs around the corner.
I don't know how to go more in depth now, but I use Claude, 4o, and o1, both mini and preview in parallel and o1 both mini and preview succeed in many tasks where the others fail.
Most things that succeed have a very clear value proposition that justifies their cost.
But I watch, and if your envisioned future arrives, that will be interesting.