Sorry, GenAI is NOT going to 10x computer programming
garymarcus.substack.com
garymarcus.substack.com
The project I am currently working on is took me about 16 hours to get to an MVP. A hobby project of similar size I did a few years ago took me about 80 hours. A lot of the work is NOT coding work that an LLM can help me with.
10x over everything is overstating it. However, if I can take my star developer who is already delivering significantly more value per dollar than my average guys and 3x him, that's quite a boost.
They can get you 95% of the way there, the question is whether fixing the last 5% will take more time than doing all of it yourself. I presume this depends on your familiarity with the domain.
I agree that the upper bound is pretty high, with more straightforward MVPs, it can generate large swabs of code for you from a description. Also for little DSLs that I haven't memorized, generating a pyproject.toml from a bunch of constraints is so much nicer/quicker than reading the docs again every few months.
Where I'm working the star developers don't have a great culture of being conscientious about writing code their non-star colleagues can read and modify without too much trouble. (In fairness, I've worked exactly one place where that wasn't the case, and that was only because avoiding this problem was a personal vendetta for the CTO.)
We gave them Copilot and it only compounded the problem. They started churning out code that's damn near unreadable at a furious pace. And that pushed the productivity of the people who are responsible for supporting and maintaining their code down toward zero.
If you're only looking at how quickly people clear tickets, it looks like Copilot was an amazing productivity boost. But that's not a full accounting of the situation. A perspective that values teamwork and sustainability over army-of-one heroism and instant gratification can easily lead to the opposite conclusion.
That one team is super productive at the cost of everyone else grinding to a halt.
You don't necessarily even have to do anything more than that. Just impose the rate limit on them and they'll automatically and unconsciously stop doing all sorts of little things - mostly various flavors of corner cutting - that make life harder for everyone around them.
FWIW, the problem you’re talking about isn’t one I’ve encountered very often at most of the companies I’ve worked at, and I’ve worked at small startups and large enterprises. I genuinely wonder why that is.
I'd argue that you're "starring" the wrong people, then.
I've known and worked with developers that are mind-boggling productive - meaning they solve problems very quickly and consistently. Usually, that's at the expense of maintainability.
I've also known and worked with developers whose raw output is much lower, but significantly higher quality. It may take them a week to solve the problem the other group can get out the door in half a day - but there are solid tests, great documentation, they effectively communicate the changes to the group, and six months from now when something breaks anyone can go in and quickly fix it because the problem is isolated and easy to adapt.
I try to be part of the second group; I'd rather get six story points done in a sprint and have them done _right_ than knock out 15 points and kicking the can down the road.
I've known exactly one developer who was in both groups - i.e., they are incredibly productive and produce top-tier output at the same time. He was 19 when I met him which would make him 27 or so now. At the time my assumption was that his pace wasn't sustainable over the long term. I should look him up again and see how he's doing these days...
A principal engineer is not someone who generates the most working code. It's someone who moves the net productivity and product impact of the whole team forward.
In my experience (mostly with Claude) generated code tends to be clean, nice looking, but ranges from somewhat flawed to totally busted. It's not ugly, messy code that works, but clean, nice code that doesn't. An inexperienced programmer would probably have no hope of fixing it up.
I was explaining it to my wife yesterday. The code is clean, but occasionally the changes are in a completely wrong file. Or sometimes it changes something that is completely irrelevant.
It's as if I had a decent programmer who occasionally loses his mind and does something non-sensical.
This is dangerous because it's easy to get lulled into a false sense of security. Most of the code looks good, but you have to be very vigilant.
It's the same with when LLMs generate writing. It sounds good, is grammatically correct, but sometimes, it's plain false.
That's what code review and mentoring are for. Don't blame tools, as a senior it's your role to pull the rest the team up. And in terms of process, it's ok to reject a PR or to delay it, business deadlines are not EVERYTHING. As a senior you're also in charge of what gets merged and when. You're the one responsible for the health of the codebase and tech debt.
On a high functioning team working on a brownfield project, the bottleneck typically isn't churning out code, it's quality control activities such as code review. The larger masses of lower-quality code that people are able to produce with Copilot make that problem worse. If the seniors are doing their jobs well then AI-generated code actually slows things down in more mature codebases, because the AI-generated code tends to take longer to review, which means that the rate of code getting into production slows down even as the rate at which developers can write it in the first place increases.
I would bet a lot of money that the people who claim huge productivity improvements with tools like Copilot are mostly self-reporting their own experience with solo and greenfield projects. That's certainly strongly implied by the bulk of positive anecdotes that include sufficient detail to be able to hazard a guess. It's also consistent with the studies we're seeing.
Probably the most annoying example of this is emitting one-off boilerplate code when an existing library function should have been used. For example AI-generated code loves to bypass facades and bulkheads, which degrades efforts to keep the system loosely coupled and maintainable.
Just yesterday I was cleaning up some Copilot-generated code that had a bunch of call-site logic that completely duplicated the functionality of an existing method on the object it was working with. I can guarantee you that I spent a lot more time figuring out what that code was doing than its author saved by using Copilot instead of learning how the class they were interacting with works.
I program daily and I use AI for it daily.
For short and simple programs AI can do it 100x faster, but it is fundamentally limited by its context size. As a program grows in complexity AI is not currently able to create or modify the codebase with success (around 2000 lines is where I found it has a barrier). I suspect it's due to exponential complexity associated with input size.
Show me an AI that can do this for a 10,000 lines complex program and I'll eat my own shorts
Now imagine 3 years from now.
From your comment, is sounds like you think that the implementation phase of LLMs is already over? And if so, how do you come to this conclusion?
What further implementations of integrating AI and programming workflows have LLMs shown to be missing?
Eventually we realized what is and isn't possible or practical to use blockchain for. It didn't really live up to all the original hype years ago, but it's still a good technology to have around.
It's possible LLMs could follow a similar pattern, but who knows.
However, if you saw the homepage of HN during blockchain peak hype, being a speculative asset / digital currency was seen almost as a side effect of the underlying technology, but it turns out that’s pretty much all it turned out to be useful for.
We shouldn’t be highly confident in any claims about where AI will be in three years, because it depends on how successful the research is. Figuring out how to apply the technology to create successful products takes time, too.
A 10x near future isn't inconceivable, but neither is one where we look back and laugh at how hyped we got at that early-20s version of language models.
It also might be that the language everyone uses 20 years from now that gives a 50X from today is just being worked on right now or won't come along for another 5 years.
The way people who would have thought that humans could never fly were not completely wrong before the airplane. After the airplane though, we are really talking about two different versions of a "human that can fly".
Not everything will grow exponentially forever
Real engineers aren’t expected to hold 10,000 lines of exact code in their head, they know the overall structure and general patterns used throughout the codebase, then look up the parts they need to make a modification.
It is one level above Claude 3.5 Sonnet, which currently is the most popular tool among my peers.
When you take into account prompt writing time and the effort of fixing its mistakes, I would have been better off doing the whole thing by hand. To make matters worse, I find that my mental model of what was created is nowhere near as strong as it would have been if I did things myself - meaning that when I go back to the code tomorrow, I might be better off just starting from scratch.
I am curious how you get to exponential complexity. The time complexity of a normal transformer is quadratic.
Or do you mean that the complexity of dealing with a codebase grows exponentially with the length?
A program of length 100 lines has the potential to be 1000x as complex as a program with length 10 lines.
Consider the amount of information that could be stored in 4 bits vs 8 bits; 2^4 vs 2^8.
As the potential to be complex grows, so does current AI's ability to effectively write and modify code at that scale.
“I would work for twenty hours,” Jackson said. “With a spreadsheet, it takes me 15 minutes.”
Not sure any other computing tool ever beat that since?
Source: Steven Levy reporting, Harper’s Bazaar 1984, reprinted in Wired: https://www.wired.com/2014/10/a-spreadsheet-way-of-knowledge...
We still don't yet have the next programming language for the ai age. I believe the biggest gains will not come from bigger models but smarter tools that can use stochastic tools like LLMs to explore the solution space and quickly verify the solution according to well defined constraints that can be expressed in the language.
When I hear about some of the massive productivity gains people ascribe to ai, I also wonder where the amazement for "rails new" has gone. We already have tools that can 100x your productivity in specific areas without also hallucinating random nonsense (though Rails is maybe not the absolute best counterexample here).
I don't have the impression GenAI can replace me, and I'm not even a 10x Dev.
Yet, it makes much of the daily task more bearable.
I suppose a reason I feel little interest in applying AI to coding may be that my experience is nothing like you describe.
That's not how you program. I've never seen anyone programing in this way unless they were just starting.
Edit: if you find yourself copying chunks and modifying them, you probably need to create an abstraction. Don't mean to be rude, just honestly that is my personal experience.
At certain very specific tasks it's going to have that 10x+ and more.
Also we've made 1000x+ gains from the times we had punch cards etc.
It will depend on the kind of projects you are working on, and the type of environment you are in. Is it a very large company or is it a side project.
In right cases it's over 10x+, in worse cases it's 10%+ maybe.
It may speak to however a type of movement that might happen is that there will be newer companies that will make sure they unleash a culture where AI first engineering is promoted, which might make it scale further than usual enterprise work currently can do. If a company can start small and lean they will be in a position to have a huge multiplier from AI. It's harder to get to that setup from a pre-existing large corporation stand point.
However, in light of Gary's other AI related writings[0], it is clear he's not saying "people, be clear headed about the promises and limits of this new technology", he's saying "genai is shit". Yes, I am attacking the character/motive instead of the argument, because there's plenty of proof about the motive.
So I am kind of agreeing (facts as laid out) and vehemently disagreeing (the underlying premise his readers will get from this) with the guy at the same time.
[0]: Here are a few headlines I grabbed from his blog:
- How much should OpenAI’s abandoned promises be worth?
- Five reasons why OpenAI’s $150B financing might fall apart
- OpenAI’s slow-motion train wreck
- Why California’s AI safety bill should (still) be signed into law - and why that won’t be nearly enough
- Why the collapse of the Generative AI bubble may be imminent
—- Take a look at this code from my multi-mode (a la vim or old terminal apps) block-based content editor.
I want to build on the keyboard interface and introduce a simple way to have simple commands with small popup. E.g. after doing "A" in "view" mode, show user a popup that expects H,T,I, or V.
Or, after pressing "P" in view mode - show a small popup that has an text input waiting for the permission role for the page.
Don't implement the changes, just think through how to extend existing code to make logic like that simple.
Remember, I like simple code, I don't like spaghetti code and many small classes/files.
It's one of those things where a language aware autocomplete will often save me more time, because it doesn't try and second guess me.
If anyone has a YT video or a detailed article, please show me!
Maybe, but they are rare. Still a welcome boost, though.
There is a significant chunk of time spent on how to address your problem, how to break it down, thinking how it will perform in real life concurrently, etc. Sitting down and coding like you are on caffeine overdrive is not the full range of “programming” (at least in my neck of the woods).
I see a lot of people mentioning 20x in here. Would they fire 20 devs tomorrow and eat their own dogfood?
At this point if we get larger context windows and LLMs become able to break down problems to include fetching relevant code into context when needed, and perhaps with the addition of some oversight training about not removing functionality, the qualty will get much, much better.
But already they are a massive time saver and they are good at some of the more boring aspects of writing code, leaving the engineer to do more systems thinking and problem solving.
Anything that requires applying things in novel ways that doesn’t have lots of examples out there already seems to be completely beyond it.
Also, it often comes up with very unoptimal and inefficient solutions even though it seems to be completely aware of better solutions when prodded.
So basically it’s a fully competent programmer lol.
We are going to be so screwed in 10 years.
The GitClear whitepaper that Marcus cites tries to account for some of these factors, but they're biased by selling their own code quality tools. Likewise, GitHub's whitepapers (and subsequent marketing) tend to study the perception of productivity and quality by developers, and other fuzzy factors like the suggestion acceptance rate -- but not bug rate or the durability of accepted suggestions over time. (I think perceived productivity and enjoyment of one's job are also important values, but they're not what these products are being sold on.)
I'm trying to understand what sort of person would end with a 3rd person anecdote like this.
"Deep learning is only good for perception" (with language one of the areas where its not good)
He really does seem like quite a good contrarian indicator if anything.
There are people who need to write a lot of emails so for those people I'm sure OpenAI will continue to deliver some kind of value but everyone else will still have to continue thinking for themselves regardless of what Sam Altman keeps promising.
> 10x-ing requires deep conceptual understanding – exactly what GenAI lacks. Writing lines of code without that understanding can only help so much.
IMHO, the real opportunity with AI, is to get developers less critical in being the only ones who translate conceptual understanding of the business problem into code and get developers more focused on the domain they should be experts in which, imho, is the large scale structure of the system and how that maps to actual computes.
What I mean by that... developers, are, on average, not good at having deep conceptual understanding of the business domain... and maybe have that understanding in the computational domain. All of the 10x developers I have ever met are mostly that because they do know both the business domain (or can learn it really quick) and the computational domain so they can quickly deliver value because they understand the problem end-to-end.
My internal model for thinking about this is we need developers thinking about building the right "intermediate representation" of a problem space and being able to manage the system at large, but we need domain experts, likely supported by AI assistants, in helping to use that IR to express the actual business problem. That is a super loose explanation... but if you have been in the industry awhile, you probably have felt the insanity of how long it can take to ship a fairly small feature because of a huge disconnect between the broader structure of the system (with microservices, async job systems, data pipelines, etc, which we can think of as the low-level representation of the sytem at large) and the high-level requirement of needing to store some data and process it according to some business specific rules.
I have no idea if we actually can get there as an industry... it is more likely we just have more terrible code... but maybe not
This tools help a bit for initial mocks, but even that, I don't like as they create code that I don't know and I need to review it all to know where things is.
When you are doing complex software, you need to build a good mental code model to know where things is, especialy when you need to start to debug issues, not knowing where things is is a mess and very annoying.
This days, I almost don't use this tools anymore, I just prefer basic line auto-completion.
Coding was only 3-5% of the time.
Also, I remember learning to race. I tried to learn to go faster through the turns.
But it turns out the way to be fast was to go a little faster on the straightaway and long turns. Trying to go faster through the tight turns doesn't help much.
If you go 1 mph faster for 10 feet in the slowest turn, it doesn't make as much difference as going .1 mph faster in a long 500 foot turn.
so...
Are we trying to haphazardly optimize for the small amount of time we code, when we should spend it elsewhere?
One of the requirements for hiring is that they be proficient with copilot, cursor, or have rolled their own llama based code assistant.
It’s like going from using manual hand tools to power drills and nail guns. Those that doubt That AI will change all of technology jobs and work in the industry are gonna find themselves doing something else.
Isn't that a bit like "I won't hire a dev unless they've written a C++ compiler"? Very dev macho, but not necessarily the best criteria for someone working on a distributed financial transactional system.
Tried spinning up a side project with just Cursor and GPT/Claude models, and we're definitely not at 'AI can build apps from just a prompt' yet. It couldn’t even create a dropdown filter like Linear’s, and the code was buggy, inefficient, and far from the full functionality I asked for. Maybe it's the niche language (Elixir), but we’re not there yet.
It certainly doesn’t write production quality code, but that entire first paragraph was science fiction a short time ago.
May be in the future, everything will become a web app because genAI became so good at it.
I guess part of the reason we can’t agree on whether they are useful is that their usefulness depends on how familiar you are with the programming task.
Where are our flying cars?
The path isnt linear.
I'm lucky to have some of the best AI integrations around, most of the time they suggest boneheadded bollocks that blocks out the proper actual autocomplete.
It is good at suggesting code hints outside of the IDE.
Any such prediction seems extremely likely to be false without a time bound, or an argument to physical impossibility.
(I personally/philosophically don't believe AGI is a natural next step for LLMs as I don't believe the English language alone which training is so heavy on encapsulates all of human ability, rather it's very honed in on English speaking countries/cultures - I also don't believe humans are very capable of creating derivative products with capabilities greater than their own - we can barely make progress on what really causes mental illness[1] - how can we claim to understand our minds so well we can replicate their functionality?)
[1]: https://www.science.org/content/blog-post/new-mode-schizophr...
However, I have never tried it on kernels, drivers, etc. Just the inverse of that :)
Now you are all mostly singing it's praises.
Good. You will be left standing when the stubborn are forced to call it quits.
What is so wrong with having a computer do amazing things for you simply by asking it?
That was the dream and the mission, wasn't it?
Unless "obfuscation is the enemy"
Most people don't want the IT guy to get into the details of how they did what they did. They just want to get on with it. Goes for the guy replacing your hard drive. Goes for the guy writing super complicated programs. IT is a commodity. A grudge purchase for most companies. But you know that, deep down. Don't you?
The bar has been lowered and raised at the same time. Amazing.