The Tower Keeps Rising
lucumr.pocoo.org
lucumr.pocoo.org
I feel like that gives an even more literal tower-rising metaphor, and that's what it feels like people using agents naively (and software engineers of lower skill or earlier-career), end up violating.
Agents are getting better at folding things into themselves, especially if you direct them to... but unfortunately I've found that the architectural instincts, even of Fable and 5.6 Sol, are still wildly behind what I reflexively achieve, say.
For sure there is an ability to have agents go back over work and try to fold it into better and better abstractions until it's sort of annealed into something good. I've done something similar on codebases that I have, but the 'high reaches' of architecture with great _prediction of how the software will evolve in the future_ in _subtle_ ways – those are, for now, out of reach of agents.
There is a part of me that wonders if it's partly just how much they can hold in their head right now, though. Even with the greatest articulation and high density of feeding them, the current setups don't allow them to hold a high-quality, sparse, 'zoomable' model of the world in their head that well yet, which we can do pretty well.
But the fact that I'm talking about it in terms of that kind of subtlety is itself promising, I guess?
We need some way to make AI-driven coding strive for parsimony.
1998 - Huge business, re-writing some vital piece of the platform in the middle of Y2K. Contract coders are expensive but also the only available people to throw at this.
The architect had mapped out the entire system down to class/method level. They'd produced a huge list of classes and methods that needed to be built. So the company hired a bunch of contract coders to build said classes and methods, including your humble protagonist. We were each given a list of methods to write up - parameters, operation, expected output. We wrote them up, and ticked them off the list. We were not briefed on how they interacted. There were no tests that we could run. There was apparently no-one checking that what we wrote in the method actually matched the spec. This was before git, so version control was extremely rough, and also before JIRA (iirc the list was an Access database).
We all realised very quickly, like the first week, that this entire project was doomed. But we were getting paid a lot of money to do this, so we just did it. It got really boring really quickly. Every day we wrote a bunch of methods, and next day got a list of the next set of methods to write. The lists just kept coming, with no idea how long the master list was, or how the classes interacted with each other, or how the system actually worked, or anything.
I left after a month. The money was good, but the boredom was driving me insane.
I learned later from friends who stayed that the whole project was canned a couple of months later when it became obvious that this was a complete waste of money and would never work.
Whenever I see a project manager staring at JIRA instead of talking to their people or looking at the codebase, I'm reminded of this project. And your comment reminded me of that ;)
A meaningful risk of course is that the tools available to the model (ripgrep + fancier semantic approaches) allow it to do a good job of reasoning over things much larger than its context window, and so it doesn't pay the penalty sufficiently to fix it.
Look at all the libraries full of books we've built. It's useful for more than mere training sets.
The limit here I think the ancestor comments are getting at is cognitive load, which is real and measured. We only have so much memory to devote to a "stack" when executing, and it's usually quite constrained.
Hence my library mention. Humans have been doing this for millennia: orienting ourselves within a library (the physical kind, full of books) and calling upon its information resources as needed to accomplish tasks (research). Ultimately, it's all just one big cache hierarchy. Your short term memory, your long term memory, the book in your hands, the desk at the library, the nearby shelves, the card catalogue, the stacks, the inter-library loan system.
To manage it all, we humans have developed our abilities for abstraction. When we build clean, tight abstractions we reduce our cognitive load. Perhaps the best abstraction we've built so far is the TCP/IP and web stack. We don't need to care at all about the hardware details of a server in order to talk to it. It's such a powerful and airtight abstraction that we take it for granted.
I'd like to hear from more people who have spent a lot of time building with LLMs, because so far what people are saying is that these models do not have the ability to reason about and build the kind of marvellous abstractions us humans have built.
I've built a lot with LLM's, my experience sort of but not really tracks that. I've had to course correct a few bad abstractions but the larger the code base becomes the better it seems to be at reusing things. Maybe this is because of types, or spec-first development (with OpenAPI), or black box integration testing - but also maybe not. But generally I have to think about the abstractions and let the LLM fill in the details with rare exception.
That said, reality at scale always come with details that will break the model, and the main roads when it happens are to ignore/reject any change proposal in the model, go in the mystic quest to reach a model that will fit it all including these new cases with an elegant simple solution, or accommodate special cases on the side until it grows too big or just percolate too fast in the main part to let it be sustainable.
Basically we've taken the "mystic quest" route, but we now have a pretty damn good data model
What's more profitable, optimizing for inference time or optimizing to increase inference time by increasing token count?
It's the infinite AI monkeys at a computer keyboard phenomenon.
Or the car on the highway that bumps left and right on the guardrails until, eventually, it arrives at its destination and nearly everybody is amazed at that great success.
The AI kool-aid drinkers are going to answer: "but that's how human code too".
And I'm really not sure about that.
That isn't to say software is perfectly built, but it's usually pragmatically built to balance costs of development and correctness - well chosen abstractions let us push up both qualities at once.
Maybe for simple one-person projects. We've long since developed methods and models to allow us to make things bigger than ourselves. Linux, SAP, etc. These software projects are not held in the mind of a single developer. But we use structure, rules, and other tools so that the pieces still fit together.
I do worry what that will mean for projects such as Linux though. Not that I think it will die or anything, but rather that it will become so fragmented that forward movement ceases.
I cannot remember who it was but there was an author who was traveling with their dog. They noticed that their dog would always pee on various tree to mark them as their territory. On their travels they ended up need some giant Red wood trees and figured, "I want to give my dog the ultimate claim of territory."
So he took his dog up to the Red wood tree and it did nothing, instead the dog wandered over to a smaller sapling and peed on that instead. The problem was the Red wood was so big and alien to the dog, it didn't recognize it as a tree.
I do wonder how many things are like that in our universe, that even if we could see them, we just wouldn't be able to understand it because it just goes beyond what we are capable of understanding. We think we have a grasp of the universe and use models to codify it but that is no guarantee that we can truly 'get it'.
Could higher level AI code be like that, would we know when we see it?
It doesn't look like we are that bright, at least collectively. And if they are individual which are really above everyone else on that matter and the rest, like maybe you but definitely not me, then their individual power seems to be unable to move us all away from our collective ill habits.
What happens if we take the most abstract libraries in any given field - and:
1. Bound to the llm to only use those as building blocks. Does it affect his reasoning ? Will it think more abstractly ?
2. Train the llm on those, so maybe it will get a feel for abstraction ?
So: better modularity.
https://blog.metaobject.com/2019/02/why-architecture-oriente...
This is the single reason why my experience with agentic development honestly kind of sucks, and wastes so much of my time.
The simplest prompt can result in the most verbose garbage ever produced, and scope-creep nobody's every seen before.
AI has a huge cognitive load problem right now. It's no wonder why so many devs say they're completely exhausted after a few hours, and not in ways they were before they picked up agentic dev.
In other words, if you can’t design a modular monolith, you can’t design a set of microservices.
I somewhat disagree:
Yes, it is possible to design a modular monolith, but thinking about the system in terms of a "minimum viable service" (but keep an eye on "viable", otherwise you can easily get into the "interaction problem") makes it much easier.
This is very similar to how you can write programs with no implicit state in a imperative programming language, but doing this in a pure functional programming language such as Haskell is much easier.
The codebases using technologies I have no idea about tend to quickly become unmaintainable and buggy, because the LLM still doesn't make good architectural choices, but the codebases that use technologies I'm familiar with basically never devolve into unmaintainability.
The difference between the two is massive, and that's why I think that a competent engineer steering an LLM in their area of expertise gets two orders of magnitude more productive, whereas someone steering an LLM in an area they know nothing about are basically producing tech debt at the speed of thought.
Shipping 100x more features per day?
I've written up my process here:
https://www.stavros.io/posts/how-i-write-software-with-llms/
The biggest thing to get right is to let the LLMs do what they're great at (code implementation from very detailed specs, and code review), and you do what humans are great (architecture and making sure the high level of the implementation is sane). That way, you get the best of both worlds, and a lot of speed at high quality.
No doubt AI boosters will claim we have found an O(1) productivity enhancer that is somehow two orders of magnitude better than anything else before. This is entering lunacy territory.
For example, I built this over a month or so and never hit a slowdown or quality issue:
Sorry, the lines have to clear what? Surely there must be some kind of constraint on "lines" that they have to overcome.
In code the thing has to become stable, can't just keep packing more and more noise onto it.
I assume one can't benchmaxx multi-year long efforts, clean architecture, taste etc as easily as these "make tests pass" tasks
Although I suspect models from Google, Facebook and Microsoft can be trained on their massive internal codebases. Whether they are is another question.
But, you would probably see a difference of scale and architecture. Larger projects that need better organization are probably more likely to be in private codebases (Linux excluded). So you might be right about the lack of private code in LLM being an issue.
At least in the past (before LLM-based code contributions got socially acceptable in some circles), in open-source project you often got very direct comments on code that was of bad quality. Yes, this was always a little bit abrasive, but it did a lot for the code quality.
For internal applications used at companies, such an abrasive behaviour is typically not accepted ("not a team player" (as if this is something bad), offended snowflakes, "not socially adept" etc.). Thus the code quality suffers quite a lot for internal applications.
Have you ever thought why salespeople have such an easy time selling some LLM for coding to big companies? Because the code quality of many internal applications is so bad, which makes even shitty LLM slop code often better than what is there.
I am pretty sure rate of improvement will be there quite long in the future — even if it will be smoke and mirrors ;)
We do have a mechanism similar to LLMs but it provides existing knowledge; it's a search mechanism, basically. New knowledge happens when we turn that mechanism off.
Maybe not exactly the same thing, but I think one related representation of this in any good codebase is documentations in the code: Good documentation always focuses on the why and not the what. i.e. things that are not represented by or representable in code. It may describe not only why a certain choice was made, but also why not something else. In general, the whole point of these documentation is to account for things that are beyond what can be understood from just reading the codebase including other existing documentation, that you may never even imagine with all the tokens in the world otherwise, which were yet a critical part of our coding choices.
If we are to believe that LLMs can build complex systems, then it's equivalent to saying that LLMs can make similar decisions equally well without any of that... which also then implies that we essentially never needed these documentation to begin with other than as redundancy for human needs.
My personal experiments show that giving them tools to access to all past sessions over a codebase helps a lot. When I coded "feature X" 2 months ago i likely specifically mentioned some constraints expressed as abstractions and, if the coding agent checks not only the code/feature it needs to implement/change but also the past sessions over that, it picks them up, and ships code that better fits the overall project.
At least, more than "architectural/design guidelines", since they are more "concrete", to the point for the task at hand.
Sessions self-preserve them for following sessions, which helps, but might also carry over stale things, so some "pruning" helps, and can be automated. Overall, as long as corrections were also made via agents and hence in sessions, they are picked up automatically.
It's not the "human zoom" you refer to, which as humans we can drive/control, but the effect seems to be similar: I read autonomous sessions where it picked up from the past the very design/architecture/abstraction points i would have driven, had I been in the loop.
In team contexts it should not be impossible to share sessions, but I not working in teams right now :D So maybe it's also a "single dev quirk".
Another point is that in a sense, learning to code goes from being told "you are using the_wrong_abstraction/this_abstraction_the_wrong_way/no_abstraction_where_you_should_have", to telling it to others/ourselves.
Since the number of knowable abstractions seems to be "at least one more than I already know", agents can actually be helpful in learning.
After a certain threshold, it basically zeros in a specific domain, but on novel domains agents taught me abstractions, when nudged towards doing that.
Great metaphor. When I hear people claim 20x productivity with AI assistance, I imagine a Tetris game where pieces fall 20x faster.
As you say, those lines still have to clear.
That probably would not be a fun game to play lol
I love this analogy, and I find it darkly hilarious that most sibling commenters don't seem to understand. Maybe you need to have worked within a million-line codebase to get it.
The only way to "clear the lines" in software is to eject them from the main codebase and into imported libraries with stable, well-documented, well-tested APIs that you very rarely if ever (security vulnerabilities?) need to touch after "stabilizing" them. Great public examples: the Go standard library, https://github.com/spf13/viper , https://github.com/uber-go/zap . Viper and zap combined are more than 20,000 lines of code (according to cloc) that I don't need to read or understand how they work - their lines have been "cleared" and all I need to know is the abstraction.
Half the joy to be found when working within massive codebases is successfully clearing lines.
It's been a few years since I read these, but if I recall the argument there, it was that Lisp makes it so easy to build stuff and scratch exactly your own itch, that there's no real strong push for lisp programmers to come together and collaborate to build non-trivial and general purpose artifacts. And that is why the landscape of public lisp software is poorer as a result, compared to languages which demand much more effort to get anything substantial done.
Armin seems to be making a very similar point about AI coding.
[1] https://www.winestockwebdesign.com/Essays/Lisp_Curse.html
It is not clear at all to me that other languages "demand much more effort" for the same end result.
It is clear that many non-lisp programmers value syntax, and many lisp programmers don't. Even many people who programmed enough lisp to have their minds blown and expanded still prefer not to program in lisp. I'm still awaiting psychological studies on this, but the rift is so large, I think there may be some significantly different brain processing going on between the two groups.
To your point, yes, it is also clear that, to the extent that lisp can match the productivity of other languages, whether it exceeds them or not, one of the tools that is needed to achieve this productivity boost in lisp is heavy usage of homoiconicity, and this results in every serious lisp program being a collection of DSLs, each of which is only understood by one person or very few people.
For me, the answer to this is economics. Even if you love lisp, there are way more companies hiring for stacks that don’t include lisp
If you then optimize for employability, which I assume most developers do (not all, but a large percentage), you might end up with not that many people practicing lisp regularly
Certainly. But why would that be? It couldn't possibly be that the programmers already there built a system using the tools of their choice, and that system is now running well and deemed maintainable enough they require additional developers for the same language, could it?
Sometimes, of course, companies change their implementation language. The most famous cases I know of this are Viaweb and reddit, where the companies moved away from lisp. The company's main motive in this, of course, is making money. If they have a good product that is deemed maintainable, why would they piss off the original team, who probably chose the language, by switching languages?
Many programmers have worked in lisp. Studies have shown that programmers routinely learn new languages and somewhat forget old ones; older programmers are typically only currently proficient in as many languages as younger ones.
A company with a successful product is going to attempt to leverage that success. Whether leveraging it looks like incrementally improving it, or like treating it as a prototype and throw it out, will typically be a long conversation between developers and management. Typically incremental improvements are greatly preferred; it is only when all stakeholders become convinced that a fresh start is required that massive changes like implementation language will take place.
The calculus, of course is changing with LLMs. See, for example, the recent migration of bun away from zig. That rewrite was spearheaded by a massively effective lone developer, the exact persona that is claimed to always prefer lisp. Sadly, he chose rust.
I can guarantee you that most of these simply didn't go far enough; they had their "minds blown" with too little. They latched onto a highlight or two, found a way of sort doing that in C++ or Javascript and moved on.
There are people who have had their "minds blown" by time/mass dilation of special relativity, or phenomena in quantum physics, who don't actually know anything about physics; they couldn't work out how far a canon ball will land fired at a 45 degree angle with a certain velocity, though they can crack jokes about Schrödinger's cat.
For those of us sticking with the Lisp program, in one way or another, it was more of a sequence of small revelations in a progression of topics. Oh, that's a good way of representing that; that is pretty nicely thought out; that fits well with that; they had that how many decades ago? Sheesh; ...
Then once you start grappling with the Lisp issues, and kick the rock farther down the road, even in some small way, you are invested.
And I can guarantee that I knew, when I wrote my comment, that I was summoning this very "No true Scotsman" argument.
> they couldn't work out how far a canon ball will land fired at a 45 degree angle with a certain velocity, though they can crack jokes about Schrödinger's cat.
It always superciliously starts with how true lisp people are the most brilliant ever.
> Then once you start grappling with the Lisp issues, and kick the rock farther down the road, even in some small way, you are invested.
And is eventually undermined by an admission of cognitive inflexibility, although not usually so quickly or in the same comment. Good job!
What made Lisp cool and powerful goes away when you do it through an LLM.
Just to be clear this is not a release, lol, I just made it to show a couple buddies. My one friend was saying he learned a lot just by reading code and pointing Claude at it and asking questions about what he didn't understand. It is provided as-is with no warranty, or installation instructions. If you really get stoked about it I could probably clean up multi user stuff a bit to get you logged on, but I think even just the idea is pretty neat.
Some have a harder time with the transition than others.
I’d look him up.
So true.
Since Nov 30, 2022 everything has become… more complex.
HTML and pre-rendering are back in, HTMx, liveview
The degaussing of CSS and the hacks we did, hell i was trying to explain how we debugged web pages in IE6 to a younger staff member today.
Some things are more complex, some things got good enough to make them less complex.
Which ones? PostgreSQL doesn't have HA in core.
FTFY
Increasing complexity is the story of mankind. It's the story of civilization.
Someone from 20,000 BC would wander around the earth trying to find food, trying not to freeze, and trying not to get eaten. Someone from 5,000 BC would be trying to grow food, hoping it rains, and hoping disease didn't wipe out the village. The second one increases the complexity from all the systems required to manage people and keep the land growing. Today the vast majority of people on earth don't grow their own food at all, and instead are busy in some way managing the complexity of a large society.
Someone from 1970-80 would think our software from pre-llm days was vastly more complex. They'd just code directly to the hardware with no abstraction layer. Now almost no one does that. We abstracted the hardware away in most cases. With cryptography libraries for the vast majority of people it's complexity is abstracted away and mostly people are told "don't try to write your own crypto because you will fuck it up".
The question now becomes, how quickly will LLMs be able to coordinate their understanding of the system they are changing?
I think the next time I see "LLMs" and "Understanding" in the same sentence, I am going to lose it....
The hierarchy is certainly higher in agriculture societies it seems, but the complexity is up for some debate
I think introducing AI to deal with this is overall a mistake though. We're just adding more complexity on top of the existing complexity. At best, it's a massive waste of hardware. At worst, we'll probably have agents introducing as many bugs as they fix as they also drown in complexity, and a lot of stuff built using these techniques are going to be fragile garbage while the overall skillset of humanity diminishes because people aren't learning the skills anymore.
Fundamentally, software does not need to be this complicated and it's a solvable problem, but it does require people that care about craftsmanship.
We should advocate hard but simple instead, but that will take more time to market, which the current trend of hot money driven market would not want, sigh. There are some good article discussing about this: https://tomauger.gitlab.io/posts/2021-06-08-simple-vs-easy/ https://samwho.dev/blog/simple-complex-easy-hard/
Back in the day we just accepted technical constraints that didn’t prevent us from doing anything mission critical. So the UIs weren’t paragons of personal expression, probably there was only one database which didn’t have any scale out features at all, and nothing scrolled infinitely. And I do think that tended to make things easier to fit into your head, because the number of distinct APIs and module interactions you needed to understand was much smaller.
But also none of it was integrated with Alexsirtanapilot, so everyday life was an endless misery.
The problem was a required input control was so heavily styled that it was genuinely difficult to recognize it as interactive. Possibly impossible for my mom, whose eyes aren’t what they used to be.
And, in place of traditional validation design language, they decided to just make the submit button invisible until the input validated.
The worst part is, someone is probably quite proud of the extra day or two they spent making sure that UI was broken.
Drowning in complexity. Paralysis of choice.
I read a comment (joke) that if you want to follow all LLM development you should have to be unemployed.
Catch-22 is it's still important to know the fundamentals so you know what to ask for, but if you don't know the esoterica, the model is eventually going to make an assumption and screw things up. And the models don't have much taste either in prose, or in coding/comment style.
Therefore it is critical that whatever AI produces is understandable to us humans. That is why we must demand that AI tools and agents produce "well designed" well-structured software. That's the bottlenexk to progress I think. Even AI can't deal with exponential complexity explosion.
It's not really news, though. Programming as Theory Building (Peter Naur) was published in the 80s, I think?
Maybe the younger entrants to this field never came across it, but even if you never came across it, it was common knowledge amongst experienced devs that understanding of the system you are about to change is crucial.
Thanks for mentioning Peter Naur’s Programming as Theory Building (1985).
I would add Fred Brooks and his The Mythical Man-Month.
The news is that Agentic Programming has made this always challenging task even more challenging.
It’s global madness fired up by continuous stream of news from LLM providers. It’s like The Verge that is almost about FAANG only, but multiplied as it is “magic” for most people, as vibe coding is “so easy” and dopamine-producing activity that it is similar to runners that don’t want to stop as it stimulates them (and in this case ofc it is healthy :)
This is so true. I am a big fan of Christopher Alexander’s “Pattern Language” concept, which addresses this exact problem! In fact he recommends developing your own pattern languages for your own domains (which of course led to the famous GoF Design Patterns book).
I have been experimenting with a “Pattern Language” skill which instructs the AI to maintain 3 pattern languages for every project. One in the business domain, one in the product domain, and one in the technical domain. It is working really well. It is always super cool to see it reference the pattern languages during planning and curate them during implementation and review.
I credit using it with keeping my 100% ai-coded projects well organized, aligned across domains, and easy to work on.
It’s mostly AI-written patterns based on my personal prompts/observations.
And here is the skill for organizing the pattern languages for a single project:
https://github.com/apinstein/skills/tree/main/skills/pattern...
I haven’t really battle tested these so hard yet but it’s a fun concept and really should get back to organizing them more intentionally. Happy to share this early phase just cause it’s cool to see the interest.
I feel these systems rising and sprawling with wee myopic agents developing out their little corners of this unknowably vast whole… a tower with 50 parapets on one side and some wacky cantilevered maiden tower on the other, and a very serviceable adobe roof over some patio for god-knows-why, and thatch over the landing next to it…
Some grotesque fatberg of designs that make sense at the level of individual design efforts, but that lack the fractal sort of levels of policy and judgment that unify the overall enterprise.
The overall language, as it were.
And language takes discipline to establish and maintain through any sufficiently large group of people—witness the company-speak or army-speak of pretty much any successful organization.
We feel like we’ve conquered the problem of talking the same language as our “Gastown Mayors” (who in turn are talking the same language as their “polecats” and so on all the way down the chain of golems)… but it’s only when it’s all built that the good Lord will humble us… that we’ll realize the understanding we thought we’d transmitted perfectly from our thrones wasn’t quite so shared as we’d imagined.
I don't know whether the author thinks this is a good or a bad thing, but in my eyes it's clearly a bad thing. Intelligence is knowing that a tomato is a fruit, wisdom is knowing not to put it in a fruit salad. AI is the the ultimate form of intelligence with zero wisdom. Actually, it's not even intelligence, it's an illusion of intelligence. If there is no human who can understand what the AI is doing it's time to stop and accept that we do not have the wisdom to contain what we are building.
Intelligence is learning to avoid using childish cliches, unless your intention was to mislead. Categorisation and understanding dependencies are hard enough problems already.
At the supermarket or taxation (Nix v. Hedden) tomatoes are a vegetable.
Interestingly, the Bible does not specify what kind of fruit was on the Tree of Knowledge. The association between the forbidden fruit and an apple was a derivative of later translations.
Imagine the consternation if the claim had become Adam and Eve ate of the forbidden tomato. Then this incessant controversy regarding the tomato's status would take on a biblical dimension... though then perhaps the Mormon claim that the Garden of Eden is to be found in the New World would actually have some credibility. Indeed, answers to some questions are not a matter of intelligence at all. Sometimes it seems to me intelligence is more about knowing what questions to ask in the first place.
That's my 2nd hardest skill when trying to use AI. 3rd hardest is better filtering.
I keep having flashbacks to learning how to use Google effectively.
I'm an arthiest, but I'm in a historically christian society so bible stuff comes up and I often delve in.
Whenever I do look into the bible, it amazes me just how extremely bullshitty vague the bible actually is. I don't understand why people argue about it, or pay attention to past arguments (many famous examples). If I wanted to argue about vague shit I'd be into philosophy or psychology or I could choose any of the many other branches of argument. Humans are not very logical. At least capitalism has a clearer scorecard.
Last nerdsniped look was about Angels due to Scott Alexander:
(The Bible describes very clearly what angels look like. [snip] This is the highest-grade antimeme I feel comfortable using as an example; if you don’t see the fnords they can’t eat you.)
Edit: Which just meta-nerdsniped me finishing up on the beaut Hispanic word "Marianismo".Though I take issue with one statement you made - "If I wanted to argue about vague shit I'd be into philosophy..." though I do sympathize.
Modern analytic philosophy in particular, if it is good, is all about trying to be exceedingly precise, often in domains where objects like the Bible (as you noted) are exceedingly vague. Philosophy has some catching up to do, but ultimately one should not disparage it so cavalierly. After all, the best philosophers were usually also the members of society most against the sort of uncritical religiosity that seems to appeal to so many, as it does now and even more severely through the ages, before the efforts of such thinkers to enlighten their fellow human were realized.
If there is a subject that seeks to answer the question of "how to ask questions" it is surely philosophy. The best mathematicians were often well acquainted. In the most famous times of "human flourishing" the philosophers and the mathematicians were one and the same.
But yes, there is a lot of bad philosophy. And there is a lot of religion pretending to be science. Sad, but it has always been true.
Mundus vult decipi, ergo decipiatur
I like that the author lets the image do the work, rather than preach at me even if I were to agree. History never repeats itself but it always rhymes.
What LLMs lack, from my experience, is two things: First, a long term memory (even a zoomed out one) and Second, the ability to execute a multi-step task without fizzling out. Upon adversity, LLMs breakdown and start making shit up. The newer models are better (and that's the difference between Fable/5.6 and GLM 5.2) but still have a very limited ability to execute any tedious planning tasks.
Don't let the agent scratch that itch - it is your hungry inner programmer expressing the need to eat i.e. do it correctly as per your taste. That morsel of work is sustenance.
It also activates the thinking gears that only kick in when I am writing the code. In the same process, the code sinks in better and becomes part of the mental model, even if I may not have paid attention to every single line.
I have also noticed that the mode switch puts me in a marginally better mood - sort of a recharge for the next bout of prompt->change->prompt.
Padmé: "For the better, right?"
Anakin: (gazes in silence)
Padmé: "For the better, right?"
For example, yesterday I came across some unit tests that didn't have error messages in their assertions. Normally, it takes me ~10 minutes to fix a handful of tests in this situation. In this case, I gave a 2-3 sentence prompt, went to the bathroom, and reviewed the result after I washed my hands. Saved me a bunch of time!
I encourage you to accept a feeling of "imposter syndrome" when using it, and keep trying new things with it. Don't feel like you have to be hands off, except when you're confident that you can be. (IE, if you think you need to spend 30+ minutes on mindless refactoring, see if you can explain it to an agent and then look at HN while it runs. You might get a good result, otherwise, it probably was time for a break anyway.)
BTW: It's important to try different models. The Claude 5.0 models are slow and give me bad results, so I'm sticking with 4.x for now.
The hard part is what text you feed it and how to judge the output.
I finally learned to let go of the code. I dont even run my C++ editor anymore.
I run frequent code and architectural reviews. Its awesome.
I don't feel the need to code review every single line of the edior I'm using. I trust it to work as promised. Same with all other tools.
So as someone managing LLMs you need to put those processes in place too. The risk is that one tries to do too much and loses overview / insight. Focusing on getting the tower tall instead of sturdy, if you will.
The code will do what you asked for in a broad sense but wherever choices arise on how to accomplish the goal it'll have made those choices incoherently.
There'll be strange validations applied inconsistently to some user inputs but not others. Data will be repeatedly sorted or converted to lower case for no reason at 3 stages. The column labels are all hard coded strings even though they are the first row of the input CSV. It'll create global functions taking data classes half the time and classes the other half.
If you look, and if you think you might one day need to update and maintain the code manually, then it's very hard to resist fixing this sort of thing, but fixing this sort of thing cancels a lot of the time saving out.
Wait what? Are there rules to vibe coding now? Are we no true Scotsmanning slop now?
https://x.com/karpathy/status/1886192184808149383?lang=en
So if you "read the diff" it isn't vibe coding, at least not by Karpathy's definition.
Obviously it's not a "rule" and you can still call it vibe coding if you want to. Maybe the meaning of the term has evolved since his tweet anyway.
Your test suite doesn’t cover all workflows. It doesn’t cover every combination of actions a user can take. So every big AI refactor while change some of those.
If this is happening frequently, your software will feel like a janky piece of unusable crap.
AIs don't care, they'll happily write 50 unit tests with slight variations and pair them with a full dockerized end to end test suite.
Now we have at least SOME tests. Are they good? Maybe, maybe not - but we have them. If one of them fails later on, we can check if it's an actual issue or the test being bad.
Which is infinitely better than having no tests because nobody can be arsed to write them or the customer isn't paying for the extra hours needed for a proper test suite.
So if the starting point is "no tests", having a bunch of "slop tests" is better, right?
Having such tests is, in fact, worse than not having tests at all. Non-existent tests are just as useful as bad tests, and are much cheaper to maintain.
it's like changing how the tower of babel should be built daily. Just because you can doesn't mean you should.
And perhaps more importantly, they don't capture more abstract properties of the code like maintainability.
AI works best in well maintained code, but unless care is taken, the AI will (today, anyway) make the code less well maintained as it goes.
If the AI is allowed to introduce mess into the codebase faster than its ability to deal with the mess increases, the codebase will eventually run into a problem that it's hard to recover from.
But you can test things and behaviours to a sufficient degree. Nobody really does 100% test coverage on real world code (outside of medical, airplanes and space of course).
Your microwave promises to heat your food at 800W for 1 minute, do you open it up and measure the magnetometer strength? Do you double-check the timer that it's accurate?
You want a program that takes in A and returns B within X milliseconds, I deliver it to you. Do you spend the time going through the whole program line by line to see how it was made or do you trust it?
Or when you buy a car do you take it to a garage and dismantle it to see whether it was correctly assembled in the factory? Or do you trust them?
Being pedantic here, but the two scenarios mentioned have regulations as well with repercussions for those products not satisfying their base requirements, whereas a lot of software is the wild-west (I won't be prosecuted for my software not actually working, in most cases).
You can give that to an agent and have it confirm the application fulfills the spec. Just like a human would have to.
And if the specific test can’t be automated, why? Is it impossible to automate or just inconvenient?
Granted there are immaterial things that we can’t automatically test yet like how the application feels to use. “Snappy” or “smooth” can’t really be quantified with code.
No, I actually want a lot more than that. I just didn't communicate all of it, because I'm accustomed to you understanding what good software is. (Or I didn't hire you in the first place, because I don't trust others to care like I do.)
If the microwave doesn't cook my food well, I can probably adjust for it; and anyway the objective power metric isn't the thing that matters to me, but the temperature and taste of the food I pull out. The code is an entirely different story, assuming I'm actually asking for code and not a finished program. Because I'm going to have to integrate it into something else.
And yeah, like the sibling comment said, the currently existing tests aren't actually the full interface. And they, too, are code, which can be wrong.
I'm not sure reading code is coming back. The ritual of reading code must come back, because that's the only way to build products that don't collapse under their own incoherence, both technically and visibly.
"just ask Claude" is fine, but it's not the end state
I guess whether AI written code means you can't build a theory of a program is an open question - but I don't see why you can't have AI written code that keeps the system knowable and extendable, in the same way that any long lived program or system can be. It's _just_ a question of being disciplined in the same way you'd need to be with developing and maintaining any long lived artefact.
The tricky part here is that you can't tell if a once-topmost part of the tower is sturdy until a great deal more tower is resting on it. Well, now a lot the economy is resting on little other than AI dreams. Your move, rational people.
I mean, yeah, hobby coding is not going away, but the feeling of exploration for me is totally gone. I don't dislike tech, I just don't see anyone being the way they used to be. People in tech are different and too many people are in tech.
Probably only part of the joy will ever be there, now. Which is weird because I did my own thing with computers til college, and don't consider myself a super sociable person. I just used to know that it was there, and now it isn't.
Is that because of the technology or because of who you were at the time?
https://www.vatican.va/content/leo-xiv/en/encyclicals/docume...
Comparing AI adoption to the Tower of Babel is a central theme to Magnifica Humanitas, but written by someone with a much deeper understanding of the Biblical narrative.
It was all good when change was slow and needed deep contemplation. Things sped up this year. Chaos is harder to defeat now. It really does need a different paradigm than "people facing into corners each doing today's task".
Not that we won't have most code written by agents. I truly beleive that to be true. But that we got so far down the "never look at code" rabbit hole that our abstractions will become so divorced from the reality that is the actual.. freaking... code. We will truly have spaghetti code, bloated towers of babble.
We'll tut-tut these times and think "if you only spent half a day thinking about the big picture, looking at code, dare I say write some code you could have saved the project a whole mess of trouble".
In the meantime we're going to be generate 10K in spaghetti what should probably be 500 reasonably understandable LoC.
Spending a day reading some code can save you $100s of tokens, and weeks of headaches.
Large Language Models are the most powerful communication tools to have ever existed - probably behind the written and spoken word and marginally behind the internet, maybe more important than phones and texting.
With understanding now on tap in this new form of infrastructure, the tower (codebase) and our coordination within it are both very customizable - up, down, sideways - per the will of the team and maintainers.
The sample size of your architectural taste, intuition, or expertise at abstraction may never be large enough for a LLM to match it. But if you want it to, it too can keep getting better and closer at this as it seeks alignment with your acceptance.
Once your tower is built, what's next?
The most interesting problems arise when you don't try and force one shared standard upon everybody yet still try and play nice.
Alternatively, power could concentrate and the winners get to decide what is valuable and not, thus cutting down the space of possible complexity by construction.
It seems to me that LLMs and particularly chatbots have already allowed for bigger scale collaboration within the LLM companies versus what was possible within the prior cohort of big platform companies.
Has the result just been taller towers, or actually a change of what is possible?
But this is just bad vibecoding? This would be bad if humans did it too. With agents or humans, you need to coordinate.
Yes, you can "add" anything you want, but if you can get that in to the main branch without a PR, it doesn't matter whether it was done by AI or human or both.
Also quite like Linus Torvalds likening of AI to a compiler.
He argues that calling AI the "author" of code is the equivalent of saying a compiler wrote your code.
He views AI as an incredibly powerful productivity layer, much like the historic progression from machine code to assemblers, and then to higher-level compilers!
Machine code was hard, but one could make it pretty efficient. Not efficiently.
Assembler was still pretty performant, for today's standard it is tip top.
So, moving on 10, 20 years from now, can someone read c++? Even html?
Today, for most devs, thats the code. We usually don't need to look at compiled output, because the code is enough. We can't just look at the prompts, because they aren't precise enough.
An AI prompt is not so precise, and an AI offers no such guarantees to respect the behaviours expressed.
This is why the primary artifact of the development process, which we review and version control, is still the code, and not the prompt.
That said, I do think there's a lot of value to be gained by recording and analysing the prompt/response loop behind the code that ends up in a codebase
We are going through a transition from a guild based software production with primitive division of labour to a machinery based one where AI is the steam engine and the job of the engineer is to build the production line, be the mechanic fixing the line, and also the assembly line worker.
So while I agree with your invocation of Marx and the alienation of labor, but I don't know if we really can compare it to the automation of manual labor.
Part of the understanding that is missed being learned by the team are the failure modes: how can this system fail and how catastrophic is it?
Now I fear only the users will learn via data loss, leaks, and hacks.
1. Perhaps with a handful of skyscrapers sprinkled in.
Why being one (I see as collaborative) was it not desired? Interpretations? Why is it seemed *more* harmful rather than good?
https://www.vatican.va/content/leo-xiv/en/encyclicals/docume...
It's odd that TFA doesn't acknowledge this.
I don't think we have hit a long enough timescale to say that we _can_ actually build on top of the AI-assisted engineering work. Tech debt tends to reveal itself later on in the process.
Personally I've totally witnessed AI-assisted work that is clearly heading towards a brick wall quite quickly. I think "mixed feelings" is my optimistic view.
People really discount how well _well engineered abstractions_ are required for AI tooling to not generate absolute garbage. And if people are relying on AI tooling for the abstractions... I think you're really tempting fate in the mid-term.
"we can, so we should".
It ended badly.
Honestly, rather than pointless debates about whether human coding is bad or AI coding is bad, I just think it's good to build tools that help me understand the world. I don't really care whether it's hand-coded or bad code.
Because most of my career has been spent as an on-site programmer. Staying at factories, visiting public institutions, deploying services for financial companies. My career is short, but I was lucky enough to work in various places.
When AI first came out, I thought I still wrote better code. But after the GPT 5 series, I've completely switched to AI and I'm now thinking about how to avoid errors and maintain larger codebases.
In the world I work in, it's common to see functions with 10,000 lines. Many people don't consider coupling or cohesion. So these days, instead of focusing on programming syntax, I'm studying programming theory and thinking about how to handle code when it becomes massive and turns into a black box through vibe coding. And I think this approach is right, because I believe I need to get used to using AI, so I keep coding with it.
But due to my cognitive limits, I've restricted myself to C# and TypeScript, which I'm comfortable with. C++ has too much to memorize and is hard to keep up with. In my region, there are very few C++ jobs, and those that exist are either extremely high-paying or garbage-tier jobs. There's nothing in between. So I stick with C# and TypeScript.
In practice, when building large programs, I often just set external configuration values I don't fully understand and code based on heuristics. I don't know the internals of Kafka, RabbitMQ, or PostgreSQL. I just know how to use them. And yet they work fine.
I feel the same way about AI code. Even if AI code is messy, if it runs, I use it. When bugs appear or performance is off, I just plug in debuggers or print statements and fix the necessary parts, like working with legacy code. Programming is so complex that if you try to understand everything, you can only design very small parts. Do the people who wrote Linux understand the entire codebase? They trust people they can rely on.
I've also reached an internal agreement to trust AI code. To support that, I'm spending time on creating rules for how to get good code from AI. Things like adding gates or CI, and seeing if that improves the code.
The problem is, I know this means no one will want to use other people's work or collaborate. The middle layer will disappear. There will be only highly admired projects or personal projects. In the past, even mid-sized projects had humans helping each other. But now mid-sized projects barely need human help. So I think projects will become increasingly polarized and become a zero-sum game.
Brooks divided complexity into two types in The Mythical Man-Month: Essential Complexity and Accidental Complexity. Personally, I think AI has greatly reduced Accidental Complexity. However, the essential difficulty, the problem of modeling, still needs to be done by humans. Because AI has no physical embodiment, it's inherently hard for it to understand domains the way humans do. Learning about something is different from experiencing it.
So I've decided to believe that vibe coding is also a valid approach. Supporters talk about compilers being deterministic, but LLMs are not deterministic. Critics say AI only produces garbage code, but I've seen that with high-quality prompts, the output becomes much better. Math PhDs say AI is good at things like theorem proving, and most of coding is similar to theorem proving.
It's not about good or bad. I've decided to believe it's just another approach. Yes, this is just a religion. My religion.
No matter how much people say vibe coding is bad, those who use it well do use it well. And there's no reason to criticize those who don't use it. I've just decided to treat this programming approach as a religion. Arguing about what's right or wrong is pointless anyway. Everyone has different values based on their environment, and convincing others is a waste of time.
People in open source communities might feel like AI code is destroying their communities. The code they used to communicate with, and the time they spent on it.
But for someone like me, who's been in delivery and on-site work, it feels like an escape hatch. It freed me from the hell of dealing with difficult people. So I've decided to rationalize it to myself: AI coding is just one way of doing things
I just think new methodologies will emerge. Instead of dividing code by functions or methods, people will think about how to divide things at a larger scale.
I'm just living to adapt to this era. I have nothing to lose anyway. I'm just waiting for the new era.
At some point the path from human language to machine code is not going to require an interstitial step writing to higher level languages as we know them.
Personally I really like using LLMs because they have made certain things a lot more productive, but I don't feel like my job has really changed at all. I have a new way of generating code, but my value has never been my (abysmal) typing speed or knowledge of algorithms or how fast I can centre a div. It's been in the stuff that LLMs still haven't pierced, and probably never will.
Typing code and using AI well are distinctly different things. I've dealt with Chinese subcontracting managers before. But managing people is different from doing code reviews and writing code. I think these two are fundamentally similar but different. Of course, it's undeniable that you need a certain level of skill to verify the output, which is why I practice typing for an hour every day. But that aside, while it's true that good coders may have a head start in using AI well, I don't think they're essentially the same thing.
The reason I say this is simple: AI has already learned more than I have in many areas. Good code often involves a lot of error handling. But humans can't cover all of those edge cases. Even something like UAF (Use-After-Free) can happen depending on timing, and pointers may not release properly. But AI catches these things surprisingly well.
So I think this way: on average, it may be true that people with higher coding skills use AI better, but I think it's hard to say that means they're better at using AI. People who are good at using AI are usually those who break functions down into more readable pieces and set boundaries better, rather than those who are good at LeetCode-style problems. That's similar to being a good coder, but it's a slightly different skill set.
I respect your opinion, but I think an additional layer will emerge.
Anyways just my opinion, I don’t mean it in an offensive way at all, I just think that people need to be careful about how they approach their usage of LLMs. Even experienced people need to be careful and ensure they don’t lose their own skills.
I agree with most of what you're saying, and I think you're giving me thoughtful advice out of genuine concern. I agree with nearly all of it.
I think of the phenomenon you're describing as "layering of input." Everything you've cautiously advised me on, I believe you're right. The atrophy of coding muscle from using AI is happening to me too. That's why I make myself type for at least an hour a day, just to maintain the coding muscle I need to actually read what the AI generates.
The problem is, I'm a freelancer, and I have to accept the reality that the market itself has already shifted around AI. Open source developers can show off the aesthetic beauty of their code to others and attract sponsors. But I'm a freelance developer, and I usually get paid per project. These days, with AI vibe coding, not only are companies handling more things in house, but deadlines have become drastically shorter compared to before. What used to be a two month project is now a one month project. And the trouble is, I charge by the month. So the workload has gone up, but my pay has effectively gone down.
I also think, as you said, that people should be more careful about how they use LLMs. But the landscape has shifted so much, and in this environment, I think the people who use it more will end up using it more adeptly. In other words, I've changed from a programmer who models with typing and UML into a programmer who models by using an LLM. I don't know which direction is the better one.
That said, I admit there's probably some self rationalization going on, because I don't really have a choice. The moment I stop working, I can't make rent. The open source world isn't really open to someone like me, an Asian guy. I upvoted you because I'm not taking your advice in bad faith. I understand where your heart is, but there's a part of me that just feels like I don't have much of a choice. Have a good day.
Where the "tower" was once a company (or team?) of human devs, it can now be a single dev and their agents.
The right engineer can likely replace non-technical co-founders with a couple LLMs. Geez, I can't wait to write that article...