LLMs are very useful tools, but if they were human, they'd be humans with sleep deprivation or early stage dementia of some kind.
All code needs to be carefully scrutinized, AI generated or not. Maybe always prefix your prompt with: "Your operations team consists of a bunch and middel aged angry Unix fans, who will call you at 3:00AM if your service fails and belittle your abilities at the next incidents review meeting.".
As for the 100% vibe coders, please let them. There's plenty of good money to be made cleaning up after them and I do love refactoring, deleting code and implementing monitoring and logging.
What it does do perfectly: convert code from one language to another. It was a fairly complex bit, and the result was flawless.
I've seen both happen. Sometimes it produced fairly good quality code on small problem domains. Sometimes it produced bad code on small problem domains.
The code is always not that great to bad at big problem domains.
Which is fine, as long as people are aware of it.
The LLM vendors are all competing on how well their models can write code, and the way they're doing that is to refine their training data - they constantly find new ways to remove poor quality code from the training data and increase the volume of high quality code.
One way they do this is by using code that passes automated tests. That's a unique characteristic of code - you can't do that for regular prose, or legal analysis or whatever.
"Even if you describe your problem (prompt) to a high standard, there is no way it can deliver a solution of the same standard."
My own experience doesn't match that. I can describe my problems to a good LLM and get back code that I would have been proud to have written myself.
If someone on my team who was a software engineer and not very junior consistently produced such low quality code I would put them on a performance improvement plan.
What the vibe-coded software usually lacks is someone (man or machine) who thought long and hard about the purpose of the code, along with extended use and testing leading to improvements.
I asked for a very, very simple bash script to test code generation abilities once. The AI got it spectacularly wrong. So wrong that it was ridiculous. Here's my reason why I think it does produce low quality code; because it does.
> "Here's a link to the commits in my GitHub repo, here's the exact prompts and models that were used that generated bad output. This exact example proves my point beyond a doubt."
I've used Claude Sonnet 4 and Google Gemini 2.5 Pro to pretty good results otherwise, with RooCode - telling it what to look for in a codebase, to come up with an implementation plan, chatting with it about the details until it fills out a proper plan (sometimes it catches edge cases that I haven't thought of), around 100-200k tokens in usually it can knock out a decent implementation for whatever I have in mind, throw in another 100-200k tokens and it has made the tests pass and also written new ones as needed.
Another 200k-400k for reading the codebase more in depth and doing refactoring (e.g. when writing Go it has a habit of doing a lot of stuff inline instead of looking at the utils package I have, less of an issue with Spring Boot Java apps for example cause there the service pattern is pretty common in the code it's been trained on I'd reckon) although adding something like AI.md or a gradually updated CODEBASE.md or indexing the whole codebase with an embedding model and storing it in Qdrant or something can help to save tokens there somewhat.
Sometimes a particular model does keep messing up, switching over to another and explaining what the first one was doing wrong can help get rid of that spiraling, other times I just have to write all the code myself anyways because I have something different in mind, sometimes stopping it in the middle of editing a file and providing additional instructions. On average, still faster than doing everything manually and sometimes overlooks obvious things, but other times finds edge cases or knows syntax I might not.
Obviously I use a far simpler workflow for one off data transformations or knocking out Bash scripts etc. Probably could save a bunch of tokens if not for RooCode system prompt, that thing was pretty long last I checked. Especially good as a second set of eyes without human pleasantries and quick turnaround (before actual human code review, when working in a team), not really nice for my wallet but oh well.
I’d say that the original claim is so context dependent that it bothers on being outright wrong.
That’s like saying that Java sucks because I’ve seen a few bad projects in Java.