Edit: Since I seem to have touched a nerve - I've been working on a project to solve this: https://www.archme.io if you want to know my thoughts on the right abstraction
Edit: Since I seem to have touched a nerve - I've been working on a project to solve this: https://www.archme.io if you want to know my thoughts on the right abstraction
I have strong disagreement because it sounds like, by analogy or proxy, we have also "solved writing"
Saying that LLMs have "reduced the cost of coding" would be boring. And using your analogy, pencils, typewriters and computers have all reduced the cost of writing, but writers are still around.
You might still need to nudge the LLM in the right direction or stop it from going off weird tangents, but none of that involves touching actual code yourself.
Ai can push a lot of keys very fast, but not always the right ones
if they need to be reminded to follow the coding standards, visible in the very code they are working on, what has been solved?
Opening the IDE and typing program code. With LLMs you don't have to use an IDE, you don't have look at program code, you don't have to care about coding standards. You ask the chatbot to write you a program/feature/fix and chatbot does it.
Is chatting with a chatbot still "coding"?
The part that isn't fully solved is just the software architecture side of things, do you want library A or library B, or write it all from scratch? LLM can do all three, but if you aren't careful it might go down a route that you don't like. But that again can be fixed with chat, "replace A with B", not coding.
How far would they get?
The problem is, without PR reviews & strict oversight, we're losing knowledge, system design & control of our codebases & products. Which is why, IMO, the coding is solved but the other parts which used to be so tightly coupled to programming are cropping up as their own issues.
But, "coding" is, except that solving it means using techniques that are still barely understood and barely described today (something I'm hoping to be able to communicate better myself about). It is now possible (at least for most consumer software, I would not say applications where human life or extreme risk is involved) to work exclusively in the domain of natural language requirements, natural language architectural/design decisions, and natural language test cases, and deliver product of equal or better software quality than average human code authors could have produced. Even in a language that you plausibly don't actually know how to code in, because the language itself can often be separable from the requirements and verification process.
If someone who doesn't know a language or framework can plausibly produce better software using it, and faster, than a team of people who do know the language/framework (and I would strongly argue that is more than plausible now with models like Astra and Fable) then it's not just "coding is less expensive."
It's like saying computers make complex calculations less expensive. That's true, but it's missing the real paradigm shift.
They haven't solved coding.
Programming is an art form. And the better you get at it, the better kinds of ideas (abstractions) you can create.
This is something today's AI cannot do.
If everyone were to permanently switch to AI for software development, software innovation would cease.
At least that's how I experience it. In the before times each non-trivial code change had a real opportunity cost as it would easily consume two days until I could even estimate whether this is worth looking deeper into.
For example, nobody on our team writes manual code anymore, we have basically set up a harness where an engineer types up the requirements for a change, the system implements it given certain constraints, we have automated unit and integration tests that are ran, and if any errors pop up, they get fed back into the loop until fixed.
But to do that, you need to actually know what you are doing - you have to have good instructions to keep the agents in check and not start making mods outside of their bounds especially when the issue is with a dependant service that is causing errors.
To solve something, there must be a defined problem, what is the problem that was solved. Or perhaps it is just "coding is solved" is the turn of phrase de jour be ause we haven't yet found a more succinct and accurate way to describe the paradigm shift
When it comes to non technical people using Ai to build things on code, the outcomes are on average pretty poor, which i see as evidence that the driver and their expertise behind the Ai matters a lot. A notable example is Terence Tao's conversation with ChatGPT, us math normies could never have done that. The same applies to coding agents ime
I think "solved coding" is taking it too far, but for many projects, the mechanical aspect of it has been removed or reduced greatly.
LLMs will have a much harder time "solving writing", because they cannot develop their own style and so are severely limited, creatively. This is less important for coding.
I still think they produce shoddy or sus code too often, an artifact of the current generation's training to try anything and everything until it "completes the task". They have a hard time even with that concept, which is part of "coding" imo
Just to make this clear: if you can define a really good PRD and sophisticated technical specs, and a strong set of tests cases to pass, at the right level of architectural granularity, plus adversarial code review processes that triangulate and weed out most mistakes, SOTA agents can write the code autonomously, at or above the quality level of most human coding teams. I call that "solved" but only if you meet those context requirements. Which is still hard, not solved, at that layer.
Solving writing is not a good analogy IMO. Writing is for human consumption, and cannot be wrapped in objective requirements and verification processes. Certain forms of writing perhaps could be (can't think of one at the moment but I don't doubt some exist), and those forms might be good analogies for being "solvable" or "solved."
I'm not typing keys, but I am very much still concerned about the quality and nature of the code. Coding to me is more than pushing keys
> if you can define a really good PRD and sophisticated technical specs
I still believe we cannot waterfall software, the idea seems like taking a step backwards. How often do we learn about an unforeseen complexity only after getting into the implementation?
In my experience with agents, it's better to be iterative and in-the-loop. Start with a decent description, have them research the code/issue, write up an initial plan/design, work iteratively on writing code and updating design doc, review and finalize the code and markdown. Then future agents will have some resources to shortcut understanding the code base.
Natural language test cases still define the expectations both at the product and architectural level and are essential for triangulating the agents on successful outcomes.
A requirements and verification approach with agents is not waterfall any more than TDD is waterfall. Does thinking ahead and doing some planning equal waterfall? Does describing how a feature works to an end user, and making some key technical decisions, before you build it, mean waterfall? Does having some sense of what you're building first mean waterfall? With agents, you can specify (with PRD and technical specs) what you THINK it should do, and in minutes or hours or at most days, have the result, which you then learn from and iterate. If you didn't fully think it through, the agents will do one of three things: 1) decide for you, which you learn from 2) stop and ask, which you learn from 3) introduce bugs, which you learn from.
It's extremely iterative.
It's one of my bigger concerns and I use this iterative approach to try and reduce it... "go look at the ./cli directory and generate me a report of inconsistencies blah blah..." or "review ./research/something.md, validate consistency, correctness, claims, ..."
... pseudo prose, have we made that a thing yet like pseudo code?
> Your concerns have moved up the stack to managing requirements, context, and verification processes.
This has always been the concern. "Add oauth to this app, there are no requirements beyond oauth working" has been an intern level task for ages. What makes software engineering hard is when the requirements start adding "well it has to use this oauth backend that isn't technically spec compliant, and the user is going to send some kind of random token you need to translate to oauth, and...". The problem isn't in writing code that will do the thing, it's figuring out exactly how that backend isn't oauth compliant and what chain of API calls I have to make to convert their random token into an oauth one, and etc.
Producing software that complies with a test suite isn't really novel. You've been able to outsource that forever. This falls apart in the same places outsourcing does; I'm sure India/Phillipines/etc/ is more than capable of iterating on code until it passes a test suite.
Today, human-language outlines / briefs / prompts are “compiled” to code which is itself then adapted to hardware. We are stretching less and less across the divide, doing less and less work on the terms of the machine. Now the farthest we’ll stretch is often formatted markdown - the most basic application of machine-parseable structure to very organic human thinking. Because we’re given the chance to be less precise, coherence suffers.
My team recently spent two weeks on a wild goose chase trying to figure out why TensorFlow Lite was generating nonsensical OpenCL kernels. Well it turns out that LLVM had a few bugs in the RISC-V assembly for our platform that was leading to silent garbage. It took combing through assembly dumps, hexdumps, a lot of pain staking debugging, and going through the TensorFlow Lite source code to to track this down.
In your opinion, if code is the wrong abstraction to be working at, how do you approach this scenario?
To your point, it's not the wrong abstraction for solving code level bugs. Just like python is not the right abstraction for solving memory corruption or pointer mis-alignments.
Which is really the same problem with coding.
The agentic model of it just taking over and doing everything is poisonous to effective long term team work.
We're well past the point where it's about the quality of the work they produce. It's the way they integrate (or rather, don't) into human practices.
This time I add another definition "when you can own it".
Except a little worse, since they were raised alone in a library, act mostly the same, and have harsh limits on personal growth.
In the first place, what does it even mean to say that code is the "wrong abstraction" for the work we do? You can't actually abstract away reality. Maybe you wish we could, that programming computers wasn't about contending with physical constraints. But it is. It will never not be important to have control over what the hardware is actually doing. Abstractions are temporary conveniences, not a replacement for understanding what is being abstracted away.
Code being the wrong abstraction is similar to assembly being the wrong abstraction if you're trying to write a web browser game. Sure it's possible to do it and yes, you may end up with more optimized code, but it's not necessary and much slower to do that. Just use Javascript.
And now, people are building larger applications & functions quicker, and it no longer makes sense time-wise to look at Javascript for-loops when understanding the work. And that's because those for-loops are generally correct, if maybe a bit under optimized. Instead, engineers need a new layer of abstraction, to understand what the code is doing without having to read each line.
LLMs write code that compiles, which they accomplish mostly by robotically attempting the task and repeatedly fixing compile errors in a loop. That's a very narrow definition of "code that works", and I would argue it is the starting point, not the finish line. My definition of code that works is more like: secure, stable, maintainable, efficient, performant, effectively bug-free, with an ergonomic interface. LLMs fail on every single fucking count. I routinely observe generated code that leaves trivial 10x or 100x gains on the table, while having severe deficiencies in every measurable approach.
Given that you talk about assembly as a bogeyman, though, I gather that you're from the generation of "software engineer" who was already writing insecure, unstable, unmaintainable, 100x inefficient, 100x non-performant, bug-ridden JS for everything. LLMs can replace this class of people, it's true.