In a more general interface they're also nice for getting a birds-eye view on a topic you're unfamiliar with.
However, just as a counterexample of how dumb they really are: I asked both Gemini 2.5 Pro and Opus 4 if there were any extra settings for VSCode's UI density and without hesitation both of them made up a bunch of 'window.density' settings.
If they can't even get something so extremely basic and well-documented right, how are you going to trust them with giving you flawless C or Typescript?
There's also a measurement vector for zero-shot LLM responses. But excelling at zero-shot is not a requirement for making LLMs useful.
The market is pointing the way, agents increase iteration capabilities, increasing usefulness. Reasoning models/architectures are another example where iterations make advances - the LLM iterates "in-band" and self-evals so that there's a better chance of a correct outcome.
All that in a mere 3.5 years since launch. To call it an autocomplete is very short sighted. Even if we reached LLMs ceiling, the choice of AI-oriented workflows (TTS, TDD, YOLO...), tooling, protocols and additional architecture adjustments (gigantic context windows, instant adaptors, speed, etc) will make up for any lack of precision the same way we work around human flaws to help us succeed in most tasks.
A human won't flip-flop from "You're right! That doesn't exist" and then straight back to "You're right, that does exist!" based on how a question is asked.
People really hold LLMs their capabilities in way, waaaay too high esteem. You have to walk a tightrope with them, and you always will unless they can fix the hallucination problem, which is quite unlikely due to how LLMs work.
* Validation: You can validate against objective signals either you or your tooling define (e.g. unit tests, compile errors, etc).
* Cost of Failure is Low: You can undo bad work, and feed errors back into the model as a signal to reduce future errors. Its not like physical domains (e.g. building a house, bridge, etc) where "undoing" is expensive and wasteful.
The models just need to be "good enough" that with enough tries the error accumulated over long jobs doesn't grow -> by adding data back into the model giving it feedback at each step you can curb this. How these agent tools sometimes achieve that is that they integrate with your build tooling stack, your IDE, your unit tests, etc etc -> they have a lot of "guard rails" to effectively curve the risk of hallucinations and/or rather when there is one to bring things back in line because they are long running processes.
TL;DR if you can't reduce the risk of bad model outputs you can mitigate the impact of the risk through retries and guard rails that feed back into the model. That's what these tools do to reduce the error rate. People are complaining about the risk of bad model output without looking at the other side - is there a way to make the effective consequence of that minimal and correct course? That's what these agent tools try to do - they want to work "like a human".
Don't get me wrong; there's a lot of signals that are hard to capture related to taste (e.g. I add a feature, it subtly changes the design of another X features already done) - and I personally find it easier to fine tune/go manual after a certain point but YMMV.
Context window size limits aside, Claude Code seems to atrophy or misinterpret existing context even prior to compactions very frequently. The more tokens within the context windows the shittier and shittier it performs—not that it's great to begin with.
Don't get me wrong; I would love to be wrong. But I do think models will get better. There's just too much money thrown at the problem and SWE seems to be the target especially I think for Anthrophic where thats their main market/use case. They don't seem to have the same diversified user base.
If I wasn't a SWE I would think there's no need to become one tbh. Just need to wait a little bit more.
Ah yes, it's just using pattern recognition from its training data to generalize abstract concepts about software so that it can apply them to my specific, complex file to find a bug that eluded multiple software engineers for a month.
It's a stochastic parrot!
Not only are they just fancy autocomplete, they are so intrusive that it increases friction for me instead of lowering it
Having to pause a second to decide if I want to tab complete a line is one thing
Having to pause a minute to evaluate an entire suggested function is jarring and completely wrecks my momentum, especially if I wind up rejecting the suggestion
And especially because if I reject the suggestion, the damn thing keeps re-suggesting stuff while I'm typing the rest out
Constant interrupting and irritating