Really does not sound like that from your description. It sounds like coaching a noob, which is a lot of work in itself.
Wasn’t there a study that said that using LLMs makes people feel more productive while they actually are not?
Really does not sound like that from your description. It sounds like coaching a noob, which is a lot of work in itself.
Wasn’t there a study that said that using LLMs makes people feel more productive while they actually are not?
I don't need a study to tell me that five projects that have been stuck in slow plodding along waiting for me to ever have time or resources for nearly ten years. But these are now nearing completion after only two months of picking up Claude Code. And with high-quality implementations that were feverdreams.
My background is academic science not professional programming though and the output quality and speed of Claude Code is vastly better than what grad students generate. But you don't trust grad student code either. The major difference here is that suggestions for improvement loop in minutes rather than weeks or months. Claude will get the science wrong, but so do grad students.
(But sure technically they are not finished yet ... but yeah)
But in all seriousness, completion is not the only metric of productivity. I could easily break it down into a mountain of subtasks that have been fully completed for the bean counters. In the meantime, the code that did not exist 2 months ago does exist.
And then quit after accepting a new job that pays them their modified value, because tech companies are particularly bad at proactive retention.
And not at all for you, because you're unlikely to retain them for long. Which makes this immaterial - AI or human, you're only going to delegate to n00bs.
(This is distinct from the question of which benefits society more, which is a separate discussion.)
But this is literally what senior engineers do most of the time? Have juniors write code with direction and review that it isn't buggy?
I mentor high-school students and watch them live write code that takes a completely bizarre path. It might technically be intentional, but that doesn’t mean it’s good or useful.
Given the nature of the statistics in question, the line between the two is extremely blurry at this point.
Some people really are going to hang on until the better end (and beyond) eh?
"AI can't code like me!" people are going to get crushed.
What does this look like in practice?
My hunches are that the to-and-fro of ideas as you discuss options or make corrections leads to a context with competing “intentions”; and that they can’t tell the difference between positive and negative experiences when it comes across each successive token in the context.
But I don’t make LMMs, so this is pure guesswork.
The rabbit hole problem is pretty rare. Usually it happens when the model flips into "stupid mode", some kind of context poisoning. If you are experienced, you know to purge the context when that happens.
In personal projects I avoid manual editing as a form of deliberate practice. At work, I only edit when it is a very small edit. I can usually explain what I want more concisely and quickly than hand editing code.
I probably would use more hand editing if I had classic refactoring tools in the IDEs similar to intellij/pycharm. Though cli based tools were a pleasant surprise once I actively started using them.
And then I waste 20 minutes bashing my head against the wall trying to write three paragraphs meticulously documenting all the key gotchas and lessons from the "good" context window with the magic combination of words needed to get Claude back in the right head space to one-shot again.
At least if I pull that off I can usually ask Claude to use it as documentation for the project CLAUDE.md with a pretty good success rate. But man it leaves a bad taste in my mouth and makes me question whether I'm actually saving time or well into sunk cost territory...
that's the issue in the argument though. it could be that those projects would also have been completed in the same time if you had simply started working on them. but honestly, if it makes you feel productive to the point you're doing more work than you would do without the drug, I'd say keep taking it. watch out for side effects and habituation though.
There are any number of things you could add to get you to any conclusion. Better to discuss what is there.
I've had the same experience of being able to finish tons of old abandoned projects with AI assistance, and I am not spending any more time than usual working on programming or design projects. It's just that the most boring things that would have taken weeks to figure out and do (instead, let me switch to the other project I have that is not like that, yet) have been reduced to hours. The parts that were tough in a creative fun way are still tough, and AI barely helps with them because it is extremely stupid, but those are the funnest, most substantive parts.
The magic of Claude is that you can simply start.
That's a significant difference. There are a lot of tasks that can be done by a n00b with some advice, especially when you can say "copy the pattern when I did this same basic thing here and here".
And there are a lot of things a n00b, or an LLM, can't do.
The study you reference was real, and I am not surprised — because accurately gauging the productivity win, or loss, obtained by using LLMs in real production coding workflows is also not junior stuff.
And if this is true, you will have to coach AI each time whereas a person should advance over time.
edit: because people are stupid, 'competitively' in this sense isn't some theoretical number pulled from an average, it's 'does this person feel better off financially working with you than others around them who don't work with you, and is is this person meeting their own personal financial goals through working with you'?
Junior is a person, not your personal assistant like LLM.
Also it is never a policy to pay competitively for the existing employees, only for the new hires.
As for humans, they might not have the motivation technical writing skill to document what they learnt. And even if they did, the next person might not have the patience to actually read it.
Also, a good few times, if it were a human doing the task, I would have said they both failed to follow the instructions and lied about it and attempted to pretend they didn’t. Luckily their lying abilities today are primitive, so it’s easy to catch.
I've been trying out Codex the last couple days and it's much more adherent and much less prone to lying and laziness. Anthropic says they're working on a significant release in Claude Code, but I'd much rather have them just revert back to the system as it was ~a month ago.
I've never had a model lie to me as much as Claude. It's insane.
This is probably because the llm is trained on millions of lines of Go with nested error checks vs a few lines of contrary instructions in the instructions file.
I keep fighting this because I want to understand my tools, not because I care that much about this one preference.
It's funny. Just yesterday I had the experience of attending a concert under the strong — yet entirely mistaken — belief that I had already been to a previous performance of the same musician. It was only on the way back from the show, talking with my partner who attended with me (and who had seen this musician live before), trying to figure out what time exactly "we" had last seen them, with me exhaustively listing out recollections that turned out to be other (confusingly similar) musicians we had seen live together... that I finally realized I had never actually been to one of this particular musician's concerts before.
I think this is precisely the "experience" of being one of these LLMs. Except that, where I had a phantom "interpolated" memory of seeing a musician I had never actually seen, these LLMs have phantom actually-interpolated memories of performing skills they have never actually themselves performed.
Coding LLMs are trained to replicate pair-programming-esque conversations between people who actually do have these skills, and are performing them... but where those conversations don't lay out the thinking involved in all the many implicit (thinking, probing, checking, recalling) micro-skills involved in actually performing those skills. Instead, all you get in such a conversation thread is the conclusion each person reaches after applying those micro-skills.
And this leads to the LLM thinking it "has" a given skill... even though it doesn't actually know anything about "how" to execute that skill, in terms of the micro-skills that are used "off-screen" to come up with the final response given in the conversation. Instead, it just comes up with a prediction for "what someone using the skill" looks like... and thinks that that means it has used the skill.
Even after a hole is poked in its use of the skill, and it realizes it made a mistake, that doesn't dissuade it from the belief that it has the given skill. Just like, even after I asked my partner about the show I recall us attending, and she told me that that was a show for a different (but similar) musician, I still thought I had gone to the show.
It took me exhausting all possibilities for times I could have seen this musician before, to get me to even hypothesize that maybe I hadn't.
And it would likely take similarly exhaustive disproof (over hundreds of exchanges) to get an LLM to truly "internalize" that it doesn't actually have a skill it believed itself to have, and so stop trying to use it. (If that meta-skill is even a thing that LLMs have ever learned from their training data — which I doubt. And even if they did, you'd be wasting 90% of a Transformer's context window on this. Maybe something that's worth keeping in mind if we ever switch back to basing our LLMs on RNNs with true runtime weight updates, though!)
These models are only going to get better and cheaper per watt.
What do you base this claim on? They have only gotten exponentially more expensive for decreasing gain so far - quite the opposite of what you say.
Humans aren’t tools.
Even if you do it by yourself, you need to do the same thinking and iterative process by yourself. You just get the code almost instantly and mostly correctly, if you are good at defining the initial specification.
The trick is knowning where the particular LLM sucks. I expect in a short amount of time there is no productivity gain but when you start to understand the limitations and strengths - holey moley.
It's more like x units of time thinking and y units of times coding, whereas I see people spend x/2 thinking, x typing the specs, y correcting the specs, and y giving up and correcting the code.
These are not _tools_ -they are like cool demos. Once you have a certain mass of functional code in place, intuition - that for myself required decades of programming to develop - kicks in and you get these spider sense tinglings ”ahh umm this does not feel right, something’s wrong”.
My advice would be don’t use LLM until you have the ”spider-sense” level intuition.
On a tangent; that study is brought up a lot. There are some issues with it, but I agree with the main takeaway to be weary of the feeling of productivity vs actual productivity.
But most of the time its brought up by AI skeptics, that conveniently gloss over the fact it's about averages.
Which, while organizationally interesting, is far less interesting than to discover what is and isn't currently possible at the tail end by the most skillful users.
Productivity is something that creates business value. In that sense an engineer who writes 10 lines of code but that code solves a $10M business problem or allows the company to sign 100 new customers may be the most productive engineer in your organization.
Taken along with the dozens of other studies that show that humans are terrible at estimating how long it will take them to complete task, you should be very skeptical when someone says an LLM makes them x% more productive.
There’s no reason to think that the most skillful LLM users are not overestimating productivity benefits as well.
I don't have to worry about managing the noob's emotions or their availability, I can tell the LLM to try 3 different approaches and it only takes a few minutes... I can get mad at it and say "fuck it I'll do this part myself", the LLM doesn't have to be reminded of our workflow or formatting (I just tell the LLM once)
I can tell it that I see a code smell and it will usually have an idea of what I'm talking about and attempt to correct, little explanation needed
The LLM can also: do tons of research in a short amount of time, traverse the codebase and answer questions for me, etc
it's a noob savant
It's no replacement for a competent person, but it's a very useful assistant
It has about a dozen or so endpoints, facilitating real time messaging.
It took me about 4 hours to build it out, fully tested with documentation and examples and readme.
About two hours were spent setting up the architecture and tests. About 45 min to an hour setting up a few of the endpoints. The rest were generated by CC. FWIW it is using layers and SRP to the max. Everything is simple and isolated, easy to test.
I think if I had a contractor or employee do this they would have coasted for a week and still fucked it up. Adding ridiculous complexity or just fucking up.
The nice thing about AI tools is you need less people. Most people are awful at their jobs, anyone can survive a few years and call themselves senior. Most teams are only successful because of the 1 or 2 guys who pull 150% while the others are barely doing 80%.
Coding at full throttle is a very intensive task that requires deep focus. There are many days that I simply don’t have that in me.
There have been many more studies showing productivity gains across a variety of tasks that preceded that one.
That study wasn't necessarily wrong about the specific methodology they had for onboarding people to use AI. But if I remember correctly it was funded by an organization that was slightly skeptical of AI.
AI coding assistant trial: UK public sector findings report: https://www.gov.uk/government/publications/ai-coding-assista... - UK government. "GDS ran a trial of AI coding assistants (AICAs) across government from November 2024 to February 2025. [...] Trial participants saved an average of 56 minutes a working day when using AICAs"
Human + AI in Accounting: Early Evidence from the Field: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5240924 - "We document significant productivity gains among AI adopters, including a 55% increase in weekly client support and a reallocation of approximately 8.5% of accountant time from routine data entry toward high-value tasks such as business communication and quality assurance."
OECD: The effects of generative AI on productivity, innovation and entrepreneurship: https://www.oecd.org/en/publications/the-effects-of-generati... - "Generative AI has proven particularly effective in automating tasks that are well-defined and have clear objectives, notably including some writing and coding tasks. It can also play a critical role for skill development and business model transformation, where it can serve as a catalyst for personalised learning and organisational efficiency gains, respectively [...] However, these potential gains are not without challenges. Trust in AI-generated outputs and a deep understanding of its limitations are crucial to leverage the potential of the technology. The reviewed experiments highlight the ongoing need for human expertise and oversight to ensure that generative AI remains a valuable tool in creative, operational and technical processes rather than a substitute for authentic human creativity and knowledge, especially in the longer term.".
> On average, users reported time savings of 56 minutes per working day [...] It is also possible that survey respondents overestimated time saved due to optimism bias.
Yet in conclusion, this self-reported figure is stated as an independently observed fact. When people without ADHD take stimulants they also self-report increased productivity, higher accuracy, and faster task completion but all objective measurements are negatively affected.
The OECD paper supports their programming-related findings with the following gems:
- A study that measures productivity by the time needed to implement a "hello world" of HTTP servers [27]
- A study that measures productivity by the number of lines of code produced [28]
- A study co-authored by Microsoft that measures productivity of Microsoft employees using Microsoft Copilot by the number of pull requests they create. Then the code is reviewed by their Microsoft coworkers and the quality of those PRs is judged by the acceptance rate of those PRs. Unbelievably, the code quality doesn't only remain the same, it goes up! [30]
- An inspirational pro-AI paper co-authored by GitHub and Microsoft that's "shining a light on the importance of AI" aimed at "managers and policy-makers". [31]
Interesting analogy, because all those studies with objective measurements are defied by US students year by year, come finals seasons.
Regardless, I'm not saying it's a cheap or practical to get high this way, especially over the long term. People probably try stimulants because folk wisdom tells them that they'll get better grades. Then they get high and they feel like a superman from the dopamine rush, so they keep using them because they think it's materially improving their grades but really they're just getting high.
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- https://arxiv.org/abs/2507.09089
""" Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%—AI tooling slowed developers down. This slowdown also contradicts predictions from experts in economics (39% shorter) and ML (38% shorter). """
Curious about specifics of this study. Because in general, how one feels is critical to productivity. It's hard to become more productive when the work is less and less rewarding. The infamous "zone" / "flow state" involves, by its very definition, feeling of increasing productivity being continuously reinforced on a minutes-by-minutes level. Etc.
actual result: 20% slower
link: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...
I hate the experience of trying to write code with them. I like to just type my thoughts into the files.
I hate trying to involve the LLM, even as a search. I want my search to feel like looking up references not having a conversation with a robot
Overall, for me, the whole experience of trying to code with LLMs is both frustrating and unrewarding
And it definitely doesn't seem more efficient or faster either
It takes an LLM 2-20 minutes to give me the next stage of output not 1-2 days (week?). As a result, I have higher context the entire time so my side of the iteration is maybe 10x faster too.
I'm a career coder and I used LLMs primarily to rapidly produce code for domains that I don't have deep experience in. Instead of spending days or weeks getting up to speed on an SDK I might need once, I have a pair programmer that doesn't check their phone or need to pick up their kids at 4:30pm.
If you don't want to use LLMs, nobody is forcing you. Burning energy trying to convince people to whom the benefits of LLMs are self-evident many times over that they are imagining things is insulting the intelligence of everyone in the conversation.
Conflating experience and instinct with knowing everything isn't just false equivalency, it's backwards.
That’s what I mean - by myself it would have taken me easily 10x longer if not worse because UI coding for me is a slog + there’s nuances about reactive coding + getting started is also a hurdle. The output of the code was still high quality because I knew when the LLM wasn’t making the choices I wanted it to make.
I feel strongly that delegation to strengths is one of the most obvious signs of experience.
Apologies for getting hung up on what might seem like trivial details, but when discussing on a text forum, word choices matter.
In other words, I don't think that you temporarily regress to "junior" just because you're working on something new. You still have a profound fundamental understanding of how technology works and what to expect in different situations.
This reminds me of the classic "what happens when you type google.com into a web browser" question, with its nearly infinite layers of abstraction from keyboard switches to rendering a document with calls to a display driver and photons hitting your visual receptors.
We might just be quibbling over terminology, however.
You’re either trusting the LLM or you still have to pay the cost of getting the experience you don’t have. So in either case you’re not going too much faster - the formers cost not being apparent until it’s much more expensive later on.
Edit: assuming you don’t struggle with typing speed, basic syntax, APIs etc. These are not significant cost reductions for experts, though they are for juniors.
Hey man, I don't bother trying to convince them because it's just going to increase my job security.
Refusing to use LLMs or thinking they're bad is just FUD and it's the same as people that prefer to use nano/vim over an IDE or it's the same as people that say "hur dur cloud is just somebody else's computer"
It's best to ignore and just leave them in the dust.