Software engineering is about managing complexity
hack8s.com
hack8s.com
However I think it's aggrandizing what human engineers actually do with remarks like "Engineers own tradeoffs." My experience is that certainly less than half of the employed software engineers don't actually give a real analysis to questions like:
"Given these constraints, this team, this business, this infrastructure, this budget, these risks, and the expected evolution of the product, what is the most appropriate way to implement X, today?"
Thus I think AI is more able to replace the average engineer more than this article admits, however the inadequacy of "average engineering" will be much more apparent now: codebases can become large/complex enough to be unwieldy in months now when it used to take 5 years [a timescale where accountability is effectively impossible].
These get overlooked so often. The way you build software if you’re at the helm vs the way you need to build it when dealing with a more/less capable team and business, especially if someone else will be doing the deployment and will need lots of consultations, is way different.
Do you think this is where it stops? This is where it begins.
Machines will be good at managing complexity too. You can't draw a line and say improvement stops here, because everything we've seen so far flies in the face of that.
I shudder to think what these models will be capable of in 24 months.
No, this is pretty much where it stops.
The models are good enough for the average coding task, and the slop they produce often is in the category of what a bad or careless dev that's being contracted out might produce.
Yea, they'll get better, but not in next-level sort of way.
The limitation is not the models or intelligence, it's the human in the loop. We're still stuck on stupid human issues ranging from usability bugs, to figuring out what the product should be, to how we should program in the first place.
I know the models are capable of sorting out issues it gets stuck on because it's writing error handling in the wrong way, or just it doesn't have the right abstractions, because we can't settle on the right way to program. I still see people arguing about dyanmic vs. static typing.
And obviously, there is a next level, but that's real singularity, and we're all out of jobs.
I consider Fable to be a super genius that was born yesterday and has memorized the internet, yet barely understands humans on a deep behavioral level.
By the end of a session, sometimes Fable and I are cooking incredibly, but then alas we have to start a fresh session. A lot of the intangibles about what are going well at that point are extremely difficult, if not impossible, to capture compactly in a “handoff” or “guidelines”.
The lived experience of working with me through the session, as represented in the current context, is what produces the higher quality results.
And likewise, the naivety it has as a newborn at the start of every session, along with the lack of deep human behavioral understanding, explains why these cutting edge models like Fable can still be so dumb in some ways while being mind-blowing in other ways.
Sure - eventually.
The question is where is the training data going to come from.
Without "training data" (the non-existent diaries of people designing complex systems and recording their thought processes) you're stuck where we are today, where the LLM is basically doing cargo-cult design and decision making - copying the outcome of decision making (problem looks like X, so I'll use design pattern Y, same as most people do), without understanding why those decisions were made.
How much of the time this matters remains to be seen, as people try to use LLMs to help design more complex bespoke software, rather than just yet another CRUD app or three.js game.
Everywhere else the percentages are quite a bit lower still (if judged by the same standards).
Prefacing that I’ve never worked at a faang. At more than half the companies I’ve worked out most of those decisions were made by non technical leadership for non technical reasons. Ranging from the reasonable(our predecessors signed a deal with Oracle a decade ago and violating it will cost us more than this project is worth) to the unreasonable(I had lunch paid for by this vendor so we’re using them now).
I also frequently ran into the problem of the process of doing that level of engineering requiring the business to make choices and being completely incapable of it. I could give 3-4 different plans with explicit tradeoffs, both in detail and with an executive summary, and what they meant for the company and even that low number of choices induced analysis paralysis in the management but they barred me from doing anything until they made the call.
Ignoring business and team scale in that question brought big damage to businesses, by developing software overly complex at times.
And the cost is of course hidden behind a salary of developers, so hard to say
The two tasks of writing code and engineering software cannot be separated without damaging the integrity of the mental model of the engineer. Having architects who didn't interact with the code always produced map/territory mismatches.
I love to say that
Some managers didn't pass the Turing testNo? Why?
Because software languages are a pretty good abstraction.
To the extent that good abstractions are in place, you can avoid looking at code specifically.
Those don't perfectly well exist, so it takes a lot of self discipline and the right tools/methods, but invariably, AI will produce better systems.
That said, its very easy to produce slop, so well see much more of it.
But mostly, it will be AI from here on in, as a matter of productivity. There are some arguments on the margins but those will fade over the next few years.
'At minimum' - the 'power tools' are here to stay.
Looking at the compiler output is a totally valid concept, but it's definitely a niche case.
It's arguably more important with Java than with the compiler output for something like C++, as C2 is much more unpredictable and dependent on runtime circumstances. You also want to be real certain that bounds and null checks are omitted as those come at a pretty big performance premium.
> No? Why?
> Because software languages are a pretty good abstraction.
No, it’s because compilers produce deterministic output. I am so tired of this argument.
If I’m not concerned with the performance of my code, I can be 100% confident that that exact code will produce the correct assembly every time. That’s why I don’t read it. Not because I don’t care.
It's entirely the nature of the abstraction.
You want it to work as expected, it does not have to produce the same thing each time.
Source: I work on an operating system.
An LLM translating a prompt to to high-level code has a much lower degree of predictability. To say an LLM prompt is a comparable abstraction is unfair, though I admit it's getting very close.
But yes, it has to fulfill some kind of contract defined by the absraction.
It's less a problem of the LLM, and more so how we use them, and the inherent tooling around it.
Programming abstractions offer interfaces to functionality that are both simplified in use and restricted in capability. (e.g. any API or compiler.) I don't see how LLMs meet that definition.
It seems more like we're talking about offloading or delegation, here. And that's a valid business tactic, certainly, but it's not a software abstraction any more than a CTO is an abstraction of a tech lead, no?
> But yes, it has to fulfill some kind of contract defined by the absraction.
I don't follow. Is the contact here the design specification for the system? If so, again, I'd argue that's not an abstraction.
That's definitely an abstraction.
IDLs are a form of abstraction, they're a requirement somewhat more formally described.
Remember UML? That was an attempt to go 1/2 layer above the code, that was an abstraction.
There were tons of tools like that.
APIs are an abstraction - maybe the best example. We write code to match exactly the behaviour defined by an APU - as long as it meets the requirement of that contract, then 'it's good'. And there could be many ways of doing that.
This is true to such an extent that I have to question the overall competence of anybody who makes the comparison. It's an enormous red flag.
I'd recommend reading Joel spolsky's leaky abstractions essay coz while it applies less and less 20 years later to things like kernel abstractions it explains very well why treating the LLM as a compiler sets you up for abject failures.
Problem with the analogy is that the strain in software engineering is necessary for an in depth understanding of the code.
The question is whether that depth of knowledge is ultimately more helpful than the speed that we can build with AI.
I realized at some point yesterday that these kinds of shortcuts would no longer be accepted
with competent LLM use.
What a load of bullshit. LLMs cut corners constantly and only handle the happy path.One example: in some old code I wrote I had assumed that the Rust Hash impl for a type is stable over time. This is not the case in general, but writing a custom hasher for a complex type is incredibly annoying, so I took that shortcut. That was fine for years, but came back to bite me this week as I was trying to update a dependency.
How much incidental complexity is due to that kind of thing? An LLM code review would flag this instantly, and one would also write a stable hash function for you.
It's hard to let that go, but you already had to in larger human organizations/collaborations where you might be assigned work on systems you never/seldom touch, or coming back to a project you haven't touched in a long time.
You don't need a mental model when you can automate the reasoning and the benchmarks that vet the reasoning. Your mental model is better spent pondering->reconsidering high level things like invariants, and then automating the the proof and implementation of those decisions.
Consider how you can just get Claude to start a workflow of 15 Fable agents to fan out over your system looking for correction/simplification/perf opportunities before spawn another wave of agents to vet the list of findings. How much time and energy and studying of the code would it have taken you to build and vet the same list?
It's not as good at system design as writing code, yet. But it feels like it's better than most of my coworkers.
I think in a few months, system architecture will have its Claude Code moment, and humans will be outclassed.
A fun little exercise you can do is design a system and write some code and then ask LLM to explain why you wrote it that way. Results are varied and interesting but in my experience rarely capture the actual why behind decisions.
There are things an engineer can do to flatten the curve - that is OP's complexity management idea - but complexity growth can never be linear as long as you are adding to the software.
I made a model/theorem for this that I posted on X: https://x.com/i/status/2027771813346820349
Code generation has exposed that verification is the central problem of software engineering. And I think it always has been.
Defining what is "correct" can be hard enough, let alone building a system that lends itself to verification, let alone spending the time to verify. Releasing software and letting users find bugs is therefore a very efficient strategy, because it spreads the burden. But you have to ride the line between losing users and getting enough feedback to find and fix the bugs that matter.
As we confront whether AI might take our jobs, I take some comfort in the idea that the world might be too complex for even the largest, best trained AI we can imagine. At a certain point, you need to simulate the whole world (or some substantial portion of it) and the cost/benefit of trying to do all that with compute may not be worth it versus using the real world (that is, humans) as your verifier.
This is actually kinda a good suggesting for everyday life as well. Many problem people face are too complex to figure out without systematic analysis, and with this method of thinking, complexity can be simplified.
Of course you have to do some swaps:
> Understand data flows -> Understand whats, hows and whys
> Choose appropriate data structures -> Choose appropriate tools
> Reason about time and space complexity -> Reason about cost and effectiveness
BTW: Is this article was about LLMs induced identity crisis? I noticed quite a few blogs written in a similar color trying to realign themselves in the new world of AI.Of course that's a reasonable thing to consider, but in addition to that, I think it's also kinda useful to remember why you started as a software engineer to begin with, what were you planning that drives you to select this path? Will AIs be a blocker of that plan, or a enhancer?
The site just goes into a reload loop on iOS?
People are shouting “yeah the hard part was never writing code, it was managing complexity” as a sort of last hurrah before AI engulfs them.
This is reality: not only can AI write code. It can manage complexity.
Prompt: read the article in this thread then execute its principles on my codebase. Write a harness and programmatic procedures that will trigger you to respond with the articles philosophy to code changes. Be vigilant and monitor every aspect constantly.
I would say for the above prompt, AI is about 60 to 70 percent as a good as a human now. A year ago it was 20 percent. The gap is closing.
Even a human could not “manage complexity” if it’s not in the right context. This is not about AI vs. human capabilities.
As an software engineer, I will readily admit that LLMs have greatly increased my output - especially on the menial work.
But now we have leadership telling everyone to "use moar AI" on everything, everywhere. I literally have observed folks dropping into incident Slack chats saying things like, "hey all - i asked Claude about this issue and then i had it write a solution. here's the PR." This feels like the kind of thing that should be a fire-able offense, but instead they're getting shout-outs from the CEO.
Hell, the next time I go on vacation, I think I could put Claude Code on YOLO mode for 2 weeks and I'd probably come back to find I'd been promoted.
I do not know how this is going to end, but I have a feeling it's going to get way darker before it gets better.
How often do engineers get a say in product direction?
Every one keeps saying that AI isnt moving the needle on the bottom line.
Well duh, code doesn't move the bottom line, features do, products do.
If you're building all the wrong things faster, all your doing is performing a speed run to a legacy code base.
I personally recommend, but I understand many people do not want to be pushed back by something they see as little more than a servant.
This setup does work to also have agents argue with each other. That can be very interesting, though you have to set them up to be very skeptical. Otherwise they will tend to read another agents assertion as authoritative off the bat.
I am convinced much of the harness/prompt engineering we are doing now will also be automated away. Within 5 years the best practices for the most popular use cases will have been found, automated and fully baked in.
Much of it is.
I think the complexity arising from 'laminar flows' etc. is just a different thing.
I wish everyone in our industry read 'No Silver Bullet, & Grug-brained Developer'.
a lot of complexity - is about what can we do now, with what we have.
It's less about writing code now but we're lying if we try to pretend it was a distraction and not a big part of the real work.
And every claim about what the job actually is or was all along has an implied (for now) at the end of it.
Which is why the productivity of people of people has never been correlated with typing speed.
In other jobs productivity is correlated with typing speed and in those jobs a typing speed like 60wpm is part of the job requirements.
It's a daily hacker news coal post regardless: "see how I can't be replaced because I do some more things". There's nothing new in it and some rehashing of this argument gets posted to HN ten times a week.
Cheers, Alberto
I've been hearing this "writing code is not what being an enginner is" mantra for years like some sort of gotcha. (It was prevalent even before AI, and I think people underestimated a lot how many people were simply incapable of writing code even given all the specs and design choices.)
Here’s an example: back in the 2000s, everyone was afraid programming was going to get outsourced to India or other countries. It didn’t happen, Sillicon Valley continues to spend billions to import engineers to work in person even though it’s 10x cheaper to hire remote outsourcers in India who are just as skilled programmers. Why do they spend 10x to move physical bodies to the office? Because it’s impossible to write good software without being in the physical context of the problem domain and team.
Similarly you cannot outsource to AI, because it cannot have complete context. No matter how good AI is, the problem is not mechanically solvable.
> To systematically investigate the role of end-user semantics of derivational traces, we set up a controlled study where we train transformer models from scratch on formally verifiable reasoning traces and the solutions they lead to. We notice that, despite gains over the solution-only baseline, models trained on entirely correct traces can still produce invalid reasoning traces even when arriving at correct solutions. More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, perform similarly to those trained on correct ones, and even generalize better on out-of-distribution tasks.
https://arxiv.org/abs/2505.13775
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens