An Honest Review of AI Programming
mropert.github.io
mropert.github.io
Underrated quote. I also found this frustrating. At one company I was at (a very old company which was trying to pivot to software engineering), we had hellish bureaucratic fights with the IT department to get access to Pycharm, Obsidian, and even GitHub. But then the AI craze dropped and management just gave us all GitHub Copilot access without us even asking.
Cory Doctorow's book The Reverse Centaur's Guide to Life After AI talks more about this. It's just a modern symptom of an age-old power struggle. The workers want more control over their craft, including quality standards and tools, but their bosses want more control over the workers.
What's happening now is bosses are feeling pressure from investors to show productivity gains from using AI, so bosses panic-push AI within their companies. Which leads to misaligned incentives like tokenmaxxing.
I've progressed in using latest Claude-kins and the GPTs as usually competent teammate/buddies, and generally know to sort out the fluff confidence with the realz (shoot, that was how I was when I was but a wee little coder lad: overconfident because of an error-free compile and one non-segfault run.)
You have to put in the time, the skill creation, the system prompt/personalization, the (sometimes adversarial) automation, the testing, verification, kicking down the loop castles (as usually caused by being cheeky with highest effort levels.)
> Unlike the silver bullets of the past (like microservices or NoSQL)
> Hallucinations are an inherent property of how LLMs work.
> It’s all marketing and buzzwords
Not a serious article or thinker. I can get this stuff on Reddit if I want to read thrice-regurgitated cliches about AI.
There is literally nothing to see in this subthread.
My agents have to pass tests meaning that if their LLM hallucinates, the agent tools capture it and not me.
Reading tests and actually catching issues requires a heck of a lot of effort. IME code reviews of tests are often more laborious than reviewing the code itself since so much if it is reasoning about corner cases.
I'm not saying don't do it, but the LOE is high. Done well I'd argue it runs close to the effort involved in just authoring those tests by hand.
And that's ignoring that the "vibe all the things" crowd is explicitly telling people not to read the code. At all.
(And if you think I'm exaggerating: an exec where I work decided to own rebuilding one of our services. They told me how recently the LLM generated multiple thousands of lines of code and they pushed it with minimal review, figuring we'll just fix the bugs as they happen...).
While technically true the hallucination rates on modern models is low and other checks can ensure that by the time a human sees it it is most likely solid.
For research there is more danger as there is less feedback loop other than other LLM scrutinising the first. For research I get it to come to a conclusion but provide me with links so I can judge. More like advanced search.
Most likely? That’s not reassuring at all. So you’re saying the other checks can result in hallucinations?
But making a decision the human reviewer disagrees with or misunderstanding a spec but maybe not asking a follow up isn't hallucination in my book.
Hallucination is just making stuff up without checking.
Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is one, it seems like it would be useful to track.
I don't particularly like these tools, but I also can't deny that for some things I'm genuinely able to get to the results I want faster with them than without. Yes, some people will use them to produce low quality software, but in my experience it's mostly people who already would have been producing low quality software as well, just at a much slower pace. Having a higher volume of low quality software is a problem, but it's distinct from the claim that the tools not ever being useful in producing high quality software.
It's possible that from reading my rebuttals you'll think I'm one of the people using these tools to produce low quality software, but it's not clear how to falsify that claim. If that's the conclusion you'd draw, then you probably weren't really open to considering rebuttals in the first place though.
Points 1-2 were clearly cited not because of their factual content, but because they boil down to the author mocking the existence of industry trends, and writing them off whole (incl. this one). This is unsound, both because it doesn't actually follow by default, and because it violates the principle of suspending (prior) judgement: https://en.wikipedia.org/wiki/Suspension_of_judgment - it sets both the author and their readers up for a specific conclusion, rather than inspiring nonbias. Hopefully one does not need to explain why this is problematic?
It is further incredibly trite to bring up how AI is an ongoing trend (self-evident), and the fact that trends distort perception (also self-evident). Them fighting fire with fire and bringing their own gut instinct is not any more intellectually respectable.
Point 3 is completely unsupported and misleading. The author probably means that hallucinations are formally unavoidable, but that on its own doesn't carry much weight, and is not the same thing.
Point 4 is factually wrong. AI is an entire academic field with a half a century of history to its name, and this is trivial to learn. Clearly not why it was cited once again either however, but because it's blatantly ill faith too.
The modern problem with AI is more the opposite, you give the AI a task that is impossible with the tools at hand and instead of saying "That doesn't work", it starts elaborate workarounds to make it happen anyway.
Might?!
Every agentic workflow should be doing this automatically for nearly a year now.
And yes, I agree with the latter. The lengths AI will go to, to get you a working solution is sometimes scary.
The times that Meta people can browbeat honest engineers are over.
Nice corporate take.
I think this observation is generally true for the kind of problems the author is working on.
But I would not make the leap to avoid using LLMs for any kind of code writing. LLMs do fantastically well in the 95%+ of the code that engineers spend time on. And for those we should leverage the technology.
It is upto us as engineers to figure out when to stop using LLMs. We are smarter than just dumping logs and half dozen specialized markdown files to a LLM and have it figure out solutions.
Also true for TV commentators, bloviating C-suiters, frequent posters on social media, politicians, etc.
Aside from that, I don't agree with the author's view on agentic workflows. Modern AI native development runs like a massive state machine, starting from MCP, local file systems, and what's usually called a harness.
I also noticed what might be a mistake in the author's domain, games. Putting aside the fact that inheritance based OOP is an outdated pattern, the suggestion to remove update() and put it into a manager class's List, then iterate with a for loop, is meant to eliminate overhead like P/Invoke costs in C#. But if Foo is still a class, a reference type, then List<Foo> is just an array of pointers scattered across heap memory. Pointer chasing can still happen. So I think that's actually bad advice.(Of course, the same issue exists in Mr. Claude's code as well.)
If the author truly wanted Data Oriented Programming(or DOD), they would have specified struct arrays or NativeArray.
This is a tricky area. The author's goal was likely to remove the per MonoBehaviour Update call cost in Unity, which is why they suggested using List<Foo>. But the instruction seems ambiguous, and I think that's part of the difficulty with AI development.
The objective function is the same, but the implementation varies and subtly differs from what I actually want.
From a design perspective, for team maintenance, the GameUpdateable abstraction might actually be better. But it's difficult. In terms of extensibility, an update manager that handles registration and expansion of multiple update targets might be over abstracting.
Writing this down makes me realize how many things I actually consider when putting code into a program. Sometimes I model how my next teammate might read it, and sometimes my words might be interpreted differently. It's really difficult.
(Or they’re working in an organization with lower budgets and not cranking the frontier models of today)
I fully agree about the cost/sustainability parts, but to suggest you can’t build a high quality coding/verifying/iterating loop for _most_ problems is disingenuous.
High performance algorithms are quite well documented so it isn't unreasonable to expect an LLM to apply them appropriately when given the ability to "see" where they need to be applied.
Unless you yourself are doing commercial game development, calling the author disingenuous puts your own pro-AI bias on display. It’s based on speculation. I prefer to take his very specific examples and experiments at face value.
I’ve done plenty of performance architecting in my day-job and rule #1 is generally “you can’t fix what you can’t see/measure”. I have a suspicion that many folks aren’t investing in letting AI actually introspect iterative execution via the appropriate harness, and are then acting surprised that it is no oracle.
Since, I'm assuming you are not a professional game developer, the parallel is speculative and sort of reaching. Therefore, is it feasible to you the author of the blog knows better than you what tools work for his chosen field and that he came to the conclusions he did in good faith?
If performance were truly critical, they wouldn't have been using Unity in the first place. They would have used Unreal, as mentioned earlier. And if they were sticking with Unity, they would have tried ECS.
Unity is fundamentally based on the template method pattern. The idea of pulling Update out and handling it in a single manager class is really more of a small scale indie game approach. It's a technique that scales very poorly.
In practice, there are many better optimization techniques for GameUpdateable. So I'm not sure why this particular example was used to demonstrate performance optimization.
Typically, you could use GameUpdateable with object pooling, which would be a safer approach. There are also many batching techniques available.
In other words, this isn't about performance. It's a technique used for small indie game development. By handling it directly through a manager, registration and removal no longer depend on the Unity framework and become manually scheduled by the user. This, in turn, means you have to handle many more edge cases, which creates additional work. This is a common pattern, sacrificing future extensibility for immediate performance gains.
It's a technique used in small indie games. Converting per frame Update callbacks into a central loop that iterates over all objects is where GameUpdateable would actually be a better choice. So rather than viewing this as an optimization for performance, it should be understood as a design choice made to make small games easier to manage.[1]
[1]https://docs.unity3d.com/Manual/events-per-frame-optimizatio...
> High performance algorithms are quite well documented
That may be true for bloom filters or what have you. But the author states the obvious: all recent games are closed-source. So any algorithms or techniques for real-time 3D that an LLM was trained on are going to be a long way behind the state of the art. The author makes that point extremely clearly, and they have credibility.
> this reads a lot like someone who decided how they feel about LLM-driven engineering ~5 months ago
TFA? No. Your comments here? Yep. Try to imagine a world where different kinds of work have different applicability of tools.
Neither you, nor the author of the article has to use LLMs for coding. But if you want to, there are some practices you should follow to get best results, and that includes setting up your environment to give the agent the best chance for success. If you have various rules that are non-obvious, you add them to your AGENTS.md file (as we have at my corp and in my little niche). The agent will then follow those rules, and learn from the surrounding code.
I'll just be blunt here - I don't believe the author would have done anything more than the bare minimum to test his pre-existing bias that coding agents are bad. I don't believe he would have put the effort into getting it writing good quality code to suit the project he was working in.
We write high performance, low latency java for trading systems. Our codebase is highly structured around this and the READMEs and AGENTs file contain the information for how to successfully write code like this, originally for human consumption, and now agents.
And it works. So, I don't trust the article.
By the same logic, you have experience in low-latency numerical decision-making, ergo you could optimise a 3D rendering pipeline.
Yeah, no.
The agent will already have seen enough super optimised code and read enough material about it. It will read the entire repo you're in, and understand how the code should be structured and written to work efficiently. And anything it doesn't get at first, you write about in the AGENTS.md file and tell it.
Then it works. If you can teach a human to write specialised 3d rendering code, you can distill the same teachings to an agent in text form, and it'll work.
I am supremely confident of this, and I bet that neither you nor the author of the article has bothered to try properly.
American Fuzzy Lop usually manages to generate valid files of any format you want, using a genetic algorithm that the author didn't even call ML. AlphaGo trained against itself.
Such things aren't impossible when you can automate the reward function, though you might have to come up with novel techniques.
I have a suspicion that a lot of the stuff people use AI coding agents for would be served just as well by, say, the WYSIWYG editor from Visual Basic 6. But we threw away that kind of technology long ago and settled for Reactslop, and now you need an LLM to write the Reactslop.
It is a common observation but I don't buy it. AI is clearly very good at Rust, but that is probably one of the least represented languages in its dataset. Anecdotally, I've also been having very good outcomes with a rather niche combination of technologies (opencv.js + JS in a browser extension) since early 2024. I would imagine there is way more C++ game code in the training set than that particular combination.
I think the more likely reason is that certain languages, projects or technologies tend to be organized in ways that are not ideal for LLMs. Specifically, I think Object Oriented approaches are not ideal for LLMs.
My theory is the key factor for effective LLM use is how effectively you can stuff the context with only the relevant data. OO tends to result in logic spread across inheritance hierarchies and templates (and even overloaded operators /shudder) which resides in a bunch of different files comingled with a whole lot of other logic. This just tends to confuse the LLM. On the other hand, I ended up using a lot more functional programming style which let me pinpoint the exact files or snippets of code relevant to a task, and the LLM pretty much never went wrong.
These days the models (and likely the harnesses) are much stronger and need much less curation of context, and hence can power through any kind of project organization. But I suspect they are still a bit sensitive to all the noise polluting their contexts and hence can produce very inconsistent results.
I've been reading this for 2 years straight. "oh you have a criticism of AI? Well they fixed that in Aeternos v Y-point-Z, which after doing all of my work also gave my wife an orgasm for the first time this year, obviously OP is using the old model".
I stopped reading when he decided to go off on some tangent about how LLMs can't do native programming or games because they weren't trained on it. Ignoring the technical issues there, an even more overt one is that, amongst a zillion other projects, even the entire source code for Unreal Engine is available and within their training corpus. Even older models were quite competent at working with Unreal and outputting idiomatic code, which is saying something if anybody's ever worked with Unreal.
I increasingly think people are writing dumb articles on purpose because it drives 'engagement' more than a straight forward and accurate post would, at least on average.
> I do admit that this approach immediately triggered my contrarian side and made me very defiant of any AI tool.
Makes sense.
> While this could be partially remedied by always asking for a primary source or citation, I dislike the idea that one has to add magical incantations to their queries to get the right results. It’s a good laugh to make fun of “make no mistake” memes, until you start having to consider similar things seriously.
Man I am genuinely stumped at the obvious lack of desire to use something in a way it's supposed to be used. LLMs are tools and like with any tool it's on us to use it properly, not hitting a screw with a hammer and saying that hammers are a very stupid tool.
You shouldn't be required to tell an information retrieval tool to actually retrieve information rather than making it up!
I mean sure, pick any analogy, still doesn't change the fact that there are right uses and wrong uses; right ways to use it and wrong ways to use it.
> information retrieval tool
How is an LLM this? It's a non-deterministic word generator.
And therein lies the problem: perception. LLMs are treated like information retrieval tools but in reality are probability machines that return plausible/mostly accurate information.
They're not searching, they're inferring and then guessing. That the guesses are often quite good means that we can easily fool ourselves with whatever it spits out. We call it hallucination but it's a feature, not a bug.
Two new types of screw, the ring-shank nonhelical fastener and the rivet, however, are changing everything. Soon, using screwdrivers will be rare except in special circumstances.
There is a middle ground between CEOisms "LLMs are literally Jesus" and people similar to this blog's author "LLMs are mostly trash."
Yeah, "part", but that's not the part I wrote about?
They're pretty flawless at any framework now, React, SwiftUI, Unity - literal skill issue if you can't get good code out of an LLM.
If an LLM can write a nuanced paragraph about any topic, it can certainly write a simple [insert popular framework] component which has considerably less potential variation.
Consider that there are 100k+ words in English, 90k+ in Spanish, and 77 in C# (38 in JavaScript!).
The LLM can write code.
Can it make software? No. Because software is a lot more than code.
But the LLM can write code.
Any anti-AI takes moving forward are going to have to acknowledge that I think