My experience is a little different. For higher abstraction languages the output is largely acceptable in my work. I always consider that LLMs don't know what I don't tell them and they have limited context to work from. Coding issues I often identify:
* Efficiency. Marginal by default. Coding efficiency problems often appear because LLMs dont usually consider the entire codebase or future plans (although they do guess at some futures). Sometimes they write/name things in ways that are lazy/wasted cycles. Most of the time, they don't.
* Security. Marginal by default. I say they do pretty good. Considering all the failure modes, not so much.
* Maintainability. Marginal by default. Mostly due to the careful consideration of modularity, upgrade paths, etc. while often taking wildly different approaches to solutions without having specific broad instructions. Even then, there can be big gaps in quality.
* Observability. Not acceptable by default. There's usually some consideration and can often one-shot.
* Portability. Not acceptable by default. Good, if you specify what those targets are. Regardless, testing validates this above the coding and models are very good at hitting functional test targets. This is less of an issue in something like Java ofc.
I do sometimes see duplicate functions, which is troubling.
Not too surprised that LLMs also don't "get it" by default?
Agents cannot be given a high level goal and then left unsupervised, for hours, without making some dumb decisions.
What people seem to be wanting is for an agent to infer vast complex data from terse simple data, which I think is probably impossible on a philosophical level. There's real information loss in language, and compute can only make guesses at the end of the day. I really don't see how we bridge that gap.
I was very optimistic about it when it was announced and saw all the demos, but a week later I find it only marginally better (and in some cases worse) than before.
Yes agents can produce code that compiles and runs, but I had to add tools to keep them on track, document their work, follow a process, check their outputs. I also use other AIs to generate developer documentation and review code.
It is like managing a bunch of idiot savant eager-to-please interns, except unlike interns, coding agents do not (yet) learn and improve on their own.
I doubt many people here are brave enough to claim their code does what is supposed to do in every conceivable case. Maybe you have high confidence in the correctness of parts of the code. Correctness of an application is murky though. Things we build are never fully correct, merely correct enough. Like maybe you're responsible for the UI in a web app and you're using your expertise to ensure it gracefully handles display across browsers and a gamut of screen sizes/form factors. But are you also verifying how it works when localized with an RtL script? Are you checking every change you make against CJK?
From time to time I try to do a pass of coalescing flows and cases and removing dead code to reduce the context and prevent the LLM from tripping over itself. But if it’s exclusively LLM maintained code I don’t care too much if there’s more of it.
Just last two weeks I had to slap Fable, three times, to stop writing 1000-2000 lines of defensive code... because of DB columns I just forgot should be NOT NULL. That was it. Nothing else. I told it that, boom, -4800 coding lines: gone.
LLMs defend the status quo and they regularly lose sight of everything bigger than the current PR they are working on.
I too am gradually making peace with the fact that LLM-maintained code does not have to be 100% readable for humans.
But this is not about readability. It's about the data model. So one concession I am willing to make is: don't care too much about the code _BUT_ manually curate the data model. So far: small wins on iteration turns and code volume producing. Too early to tell but for now I am happy with the results.
Its documentation about what the code does not do could fill whole books.
UI copy being full of slop explaining what the software does not do is another problem.
I am not convinced that a lack of negative test cases is an issue.
I do agree it generates too much code most of the time.
And the problem isn't just that it says what the software doesn't do, most of the things it claims are in fact meaningless, it's not even clearly describing something the software shouldn't do.
Really doubt we are near that being solved with non-technical folks + LLMs. I'm seeing people gleefully rebuilding products with the exact same blind spots in their understanding/logic using LLMs. Claude, etc are not seemingly able to "AGI" around goofy asks. The CSS looks a little nicer than their legacy products though, lol.
This reminds me a bit of a PhD Comics webcomic that confidently claimed we would "never" cure cancer, on the grounds that cancer is not one thing. And I don't know that we will ever actually cure cancer, but that wouldn't be the reason. Correctly noting the problem space is bigger than a layperson would initially appreciate is a lot of things, most of them helpful, but the one thing it's not is a formal a demonstration of optimizing against the problem space as a whole.
The refactor ended up adding 22,000 loc.
I went in there and quickly read through it, laughed my ass off. Reverted the work tree. Micromanaged a new refactor. Net lines of code for something really elegant and easy to reason about was -3k loc in the project.
In case you are wondering why vibe coders are doing 30k loc a day, this is why.
Better at writing code within a huge system, definitely not. Maybe in the future, but as of Astra, Fable 5.1, the answer is still no.
If anything, it works more reliably today with the smarter models.
That is never worth it. You're ruining the software you work on when you do this.
The speedup of slop production being “worth it” is what we, as a society, are having trouble evaluating at this point in time. In all likelihood it’s worth it only in the short term.
It's just that many (I guess that includes me? :D) assumed that they are better than the actually were.
AI I've used isn't a better coder than I am - it's just got a lot more hours in an hour than I do.
I see this comparison a lot, and I think it's a trap, because it invites us to confuse scaling duplicates with scaling design changes.
Duplicative mass-production was always core to software from the moment it first became "soft". A factory churning out 10,000 copies of the same book maps to 10,000 downloads of a single software release. The paper and bindings of the book may be below hand-crafted standards, but the words are largely unaffected.
In contrast, LLM-coding is the design and prototyping stage. So if we want to learn from textiles/electronics, we shouldn't be thinking of acres of looms, but instead about fashion-design, custom tailoring, determining patterns for clothes, designing new appliances, choosing circuit layouts, etc.
This is just a hypothetical example, I'm not saying that this is how it would necessarily go in all cases.
"Does that code work?" "yes" "so lets ship it" "but its not good" "but it works right?"
It's astounding to me that people can see coding get solved and not think every single one of these tasks won't be solved too.
Why do you not think these things aren't going to be completely automated? What makes these tasks special?
Fable and Astra can one-shot video games with compelling novel game loops. They can do systems programming, distributed systems, robotics. I haven't found a weak point.
Seedance 2.5 can make video better than the manual labor of VFX artists, 3D artists, and animators.
Nano Banana and GPT Image can do a better job than graphics designers.
LLMs just solved a Millennium Prize Problem, and there are probably more that will fall in the coming weeks.
Just wait. All of these things will be solved.
There is no "stopping point".
Edit:
Don't anticipate that 2036 will look anything like 2026.
Will Smith spaghetti doesn't stay that way forever. Trillions of dollars will be spent on solving these problems. They will be solved.
If you value humans intrinsically, this is necessarily the loop that will converge. I don't think humans have deep intensional a priori knowledge of the structure of reality. If we did, then we wouldn't need tools like AI because we'd be a superset of that. We can only observe and judge.
If we don't value humans, then sure, I think AI is at the point where it can kill all humans (conditional on sentience and resources etc). Two ways to solve a problem - solve the problem, or eliminate the problem statement. Plenty of easier vectors to eliminate the "problem statement", than say, try to solve problems such as making human life better. If you do value the latter though, there will necessarily be human judgers. That's how it works.
Have you tried one-shotting real distributed systems problems? What was the result and how did you verify correctness?
But lets be fair, if an expert would use AI today to build something with this, I would feel a lot more confident than not doing this.
I would start with the base architecture and add all the guardrails for a distributed system, i might even go so far to leverage the math skills of a frontier model like fable or astra. I would for sure have the proper budget for using Fable/Astra.
We haven't seen that. Maybe when we do, we will start to believe that other things will get solved.
Tell that to the mountain of failed AI slop games on Steam! As a game dev, building compelling, fun games is not even something humans are good at doing consistently. The AI can build the tech, but it can't make something 'fun' yet (unless your bar for fun is simply that a tool created a thing).
Will all these things get automation? Yeah sure. But the idea that they will be perfect automated solutions applicable in all cases is just marketing, it’s not reality.