However, most of the engineers I respect have gone from being skeptics a year ago to convinced today. I don’t personally know any true holdouts any more. If there are studies that disprove productivity gains more than six months ago, I’m happy to believe that it was true of the AIs that were available at the time. But I’m going to need something much more recent before I disbelieve my lyin’ eyes where it pertains to the AIs available today.
Here is the report:
https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
And my commentary:
My view is that it's not really about how good the models are - it's about how we're using them. Understanding what you've built is an important part of value creation, and LLMs eliminate that.
I currently don't have work access to Claude Code, but most of my teammates do. Watching from the outside, the cycle seems to look like this:
1. Experience some success, which hooks you into relying on AI.
2. The AI keeps failing at some task, but you don't want to stop. Keep trying over and over again.
3. Run out of tokens and take a break.
Now, sometimes 1 doesn't happen. Sometimes 2 doesn't happen. 3 is a certainty though.
Now, if you told me that the productivity gain from 1 is enough to offset the loss from 2 and 3, I could believe you. But I also wouldn't be surprised if it didn't.
Even if you take for granted that AI is as good as the best people say in writing code. And Ive spent a lot of time generating codes, I won’t disagree - Then the question becomes - does this change your daily incentives such that you reach for code as the solution to your problems rather than something else (coordinating with your colleagues? Product management? Planning and Design?
So from a holistic perspective, I think intentionally limiting your own AI usage is the best approach for maximum long-term productivity.
But what if the problem you’re trying to solve is the altogether too often problem of like getting teams that are dependent on you to upgrade the library they use. And what if the library is a breaking change, and last year they upgraded to the library on your advice and it broke production and now they’re suss and want to accept all changes, and integrating that library change isn’t in their critical path so they’re just not going to spend time on it, even if you submit the MR them. Even if you show them their tests pass after the change.
Importantly to the above, you probably need more devs to do more of the above in parallel. You don’t hire devs to write more code, you hire more devs to carry on the mental load of a broader scope of work. Even in the before times, so much code got stuck at the integration step.
But because all that is hard, instead you go and codegen to fix an obscure bug that sure makes a few customers happy, but no one thought was a limiting factor for paying your company more money.
It’s not that I don’t think AI can help, I think it’s a prerequisite for the job and everyone should use it. It’s more that I think in the grand scheme of things, people will bias towards using it for tasks that aren’t in the critical path - refactors, tech debt, bug smashing, tool building; and I think it could really help devex and that’s good.
But I think people are bad at knowing the difference between “my job feels a bit easier” or “I’m more productive” and “this task had an impact on the bottom line” and when you extrapolate that out to a whole engineering org, that’s where the productivity statistics get lost.
I’ll addd one data point to this is like this thread itself. So many people on AI skepticism threads point to their own subjective experience as evidence we’re not in a bubble, and sort of ignore the entire concept of economics. I’m not saying we’re in 100% in a bubble, but subjective experience isn’t great evidence of it.
And this is just sort of one of the factors, what about the increased cost and mental load of supporting more software? What about junior engineers who feel pressured to ship work but don’t actually learn the software engineering? What about lost context from not intimately understanding your software?
Although if this theory is true — that AI helps with coding but coding is not the friction point in organizations with multiple humans, even that should allow faster iteration by allowing one human to do more coding therefore reducing the size of teams required to make some programs. You should see good acceleration in solo shops too.
I’m a platform engineer. The primary failure mode for platform engineers is building tools people don’t want. AI doesn’t really make that easier. Or it can but it can also make it harder by making it easier to chase down ideas that you don’t get traction on. And I think that - net balanced across the organization is probably why productivity gains get sort of averaged out.
For sure I think solo devs who have a system are seeing gains, as long as you can I think have the discipline to have a process that includes feedback and learning and your not just feeding off of dopamine hits of one shotting features but yeah. I mean for solo devs the code was never really the limiting factor, it was product-market fit and marketing.
So solo devs who have a system may be laughing themselves all the way to the bank, but we may not see a lot of net new solo devs.
But if code is cheap now then it’s sort of inherently devalued. 2 8 person startups can probably relatively easily find a dev with AI experience to rocket ship their code generation, which means the basic skills of talking to customers, change management, and building the right thing become even more valuable.
Even solo devs I wonder - almost every post-mortem of a failed company goes “I wish we had spent more time talking to customers and less time writing code”
Again if you can get the discipline right, maybe as a solo devs you can get more work done faster and spend more time with your family. That’s incredibly valuable!
But if you go and add a big new feature, or a second product - unless your community is primed for constant growth(no man’s sky is one community where more more more seems good) you’re just growing the surface area where all the other skills are more necessary.
I think this is right. They are much better applied as editors than authors, IMO.
The key thing is stay in control of your output. i.e. understand it thoroguhly. I think you let the LLM make decisions you don't really understand, you're increasing the likelihood of introducing defects that are expensive to address.
The thing people I think have a hard time seeing is that "I go faster" does not mean "more features get finished".
It's a scale issue, and one scale is better than the other. People only pay for finished features, they do not pay for how much code you emit.
In my field - operations - productivity is usually described as some rate of production for a specific asset. 100 widgets / machine / hour - for example.
"My productivity is 3 PRs / day with the LLM as opposed to 1 PR per every three days". That's how I think people are thinking about it.
My point is that's not the same thing as value. I.e. what people will pay for.
“This random part of the code is slow, I used an LLM to generate a PR that speeds it up.”
Okay, you optimized the part that’s not a bottleneck, sped up nothing and cost the company $100 in tokens. Good job?
"If an LLM builds a feature, and no one uses it, did it make value?"
You're right my analysis is at variance to what Faros.ai says. I think they interpret their data trying to rescue utility for the dominant patterns of LLM use.
But I think to anyone who is experienced with process improvement or queuing theory, their interpretation is clearly weak. Rework is a huge problem in queue systems, and they mostly just elide the throughput impact of an 860% increase in code churn coupled to a massive spike in bugs.
Obviously draw your own conclusions. But I don't think because I disagree with the interpretation of the people who originated the data makes me wrong.
EDIT: In fact, parent comment has a link to some numbers.
[EDIT: Most] people don't want to go through the numbers. Ok. But there's a history here. When people don't want to see the numbers, certain kinds of things tend to happen.
Code acceleration is great, but.... something precedes that. Vision and strategy re. expansion of offerings and businesses. Once a firm reaches maturity in what it offers and is only touching the edges - this code acceleration is literally useless when you factor in all of the trade-offs.
This is a good thing - it means fat and slow incumbents are sitting ducks to be out-witted by creative and imaginative founders, which is healthy for a well-functioning economy.
Now the economics of existing frontier models are not sustainable - its looking like a mix of the airline (supersonic vs subsonic) and EV industry with China in the background providing decent offerings at much lower prices.
I admit that if a small team or an individual uses an LLM, it's likely they can create value faster.
I think as soon as you don't own the responsibility for the defects you generate with an LLM, their use starts to destroy value. Regardless of product maturity.
This is what I think the data says.
I actually think this is precisely the reason LLMs can't be the basis for a technological revolution. Because it's only one way.
Like, if you have a compiler, and it has a bug. You can discover if that bug is influencing your code execution and patch it. You can go both up and down the stack.
With LLMs, there is no way to patch it's translation function. You have to rely on it to forward process.
I don't think there is any way to avoid us understanding our tech stacks.
If you are producing something that delivers a far better experience, irrespective of what's under the hood (see Claude Code et al), you will decimate an incumbent who is trying to use LLMs in the context of incrementally improving a mature product.
LLMs are suited for the development of revolutionary innovation, not incremental.
I think I just disagree about the power of the LLM to deliver revolutionary innovation. That's something you do. Not the machine.
And, pretty soon on your journey to scale, the LLM becomes a hinderance rather than a help.