Developers notoriously overestimate the productivity gains of AI, especially because it's akin to gambling every time you make a prompt, hoping for the AI's output to work.
I'd be shocked if the developer wasn't actually less productive.
Developers notoriously overestimate the productivity gains of AI, especially because it's akin to gambling every time you make a prompt, hoping for the AI's output to work.
I'd be shocked if the developer wasn't actually less productive.
The baseline isn't what it would have taken had I set aside time to do it.[1] The baseline is reality. I'm easily getting 10x more projects done than in the past.
For work, I totally agree with you.
[1] Although it's often true even in this case. My first such project was done in 15 minutes. Conceptually it was an easy project. Had I known all the libraries, etc out would have taken about an hour. But I didn't, and the research alone would have taken hours.
And most of the knowledge acquired from that research would likely be useless.
[0] tasks included making games from scratch and resolving bugs we put into template projects. There's no perfect tasks to test on, but this seemed sufficient
Obviously you cannot generalize that to all software development though.
This is in part because context limits of large code bases and because the knowledge becomes more specialized and the LLM has no training on that kind of code.
But people are making it work, it just isn't as black and white.
Just today on HN I’ve seen claims of 25x and 10x and 2x productivity gains. But none of it starting with well calibrated estimations using quantifiable outcomes, consistent teams, whole lifecycle evaluation, and apples to apples work.
In my own extensive use of LLMs I’m reminded of mouse versus command line testing around file navigation. Objective facts and subjective reporting don’t always line up, people feel empowered and productive while ‘doing’ and don’t like ‘hunting’ while uncertain… but our sense of the activity and measurable output aren’t the same.
I’m left wondering why a 2x Microsoft of OpenAI would ever sell their competitive advantage to others. There’s infinite money to be made exploiting such a tech, but instead we see highschool homework, script gen, and demo ware that is already just a few searches away and downloadable.
LLMs are in essence copy and pasting existing work while hopping over uncomfortable copyright and attribution qualms so devs feel like ‘product managers’ and not charlatans. Is that fundamentally faster than a healthy stack overflow and non-enshittened Google? Over a product lifecycle? … ‘sometimes, kinda’ in the absence of clear obvious next-gen production feels like we’re expecting a horse with a wagon seat built in to win a Formula 1 race.
I'm currently using AI (Claude Code) to write a new Lojban parser in Haskell from scratch, which is hardly something "super basic and common". It works pretty well in practice, so I don't think that assertion is valid anymore. There are certainly differences between different tasks in terms of what works better with coding agents, but it's not as simple as "super basic".
Some of the AI generated I've seen has been decent quality, but almost all of it is much more verbose or just greater in quantity than hand written code is/would be. And that's almost always what you don't want for maintenance...
However if the difference is between doing a project vs not doing is, then the gain is much more than 10x.
Last month:
128 files changed, 39663 insertions(+), 4439 deletions(-)
Range: 8eb4f6a..HEAD
Non-merge commits: 174
Date range (non-merge): 2025-12-04 → 2026-01-04 (UTC)
Active days (non-merge): 30
Last 7 days: 59 files changed, 19412 insertions(+), 857 deletions(-)
Range: c8df64e..HEAD
Non-merge commits: 67
Date range (non-merge): 2025-12-28 → 2026-01-04 (UTC)
Active days (non-merge): 8
This has a lot of non-trivial stuff in it. In fact, I'm just about done with all of the difficult features that had built up over the past couple years.One could argue that some humans write that way, but ultimately it does not matter if the text was generated by an LLM, reworded by a human in a semi-closed loop or organically produced by human. The patterns indicate that the text is just a regurgitation of buzzwords and it's even worse if an LLM-like text was produced organically.
Recently I’ve been using it to write some async rust and it just shits the bed. It regularly codes the select! drop issue or otherwise completely fails to handle waiting on multiple things. My prompts have gotten quite sweary lately. It is probably 1x or worse. However, I am going to try formulating a pattern with examples to stuff in its context and we’ll see. I view the situation as a problem to be overcome, not an insurmountable failure. There may be places where an AI just can’t get it right: I wouldn’t trust it to write the clever bit tricks I’m doing elsewhere. But even there, it writes (most of) the tests and the docs.
On the whole, I’m having far more fun with AI, and I am at least 2x as productive, on average.
Consider that you might be stuck in a local (very bad) maximum. They certainly exist, as I’ve discovered. Try some side projects, something that has lots of existing examples in the training set. If you wanted to start a Formula 1 team, you’re going to need to know how to design a car, but there’s also a shit ton of logistics - like getting the car to the track - that an AI could just handle for you. Find boring but vital work the AI can do because, in my experience, that’s 90% of the work.
I've started and finished way more small projects i was too lazy to start without AI. So infinitely more productive?
Though I've definitely wasted some time not liking what AI generated and started a new chat.
Yes that's already been well established.
I gave agents a solid go and I didn't feel more productive, just became more stupid.
Perhaps you're not playing to their strengths, or just haven't cracked the code for how to prompt them effectively? Prompt engineering is an art, and slight changes to prompts can make a big difference in the resulting code.
- Providing boilerplate/template code for common use cases
- Explaining what code is doing and how it works
- Refactoring/updating code when given specific requirements
- Providing alternative ways of doing things that you might not have thought of yourself
YMMV; every project is different so you might not have occasion to use all of these at the same time.Your list gives me a starting point and I'm sure it can even be expanded. I do use LLMs the way you suggested and find them pretty useful most of the time - in chat mode. However, when using them in "agent mode" I find them far less useful.
I agree 10x is a very large number and it's almost certainly smaller—maybe 1.5x would be reasonable. But really? You would be shocked if it was above 1.0x? This kind of comment always strikes me as so infantilizing and rude, to suggest that all these developers are actually slower with AI, but apparently completely oblivious to it and only you know better.
Maybe shocked is the wrong term. Surprised, perhaps.
So far it’s just crickets.