The great AI delusion is falling apart
mikemcbrideonline.com
mikemcbrideonline.com
It's not a bad paper, but it's also turning into a fantastic illustration of how much thirst there is out there for anything that shows that AI productivity doesn't work.
I just learned there was a 4 minute TV news segment about it on CNBC! https://www.youtube.com/watch?v=WP4Ird7jZoA
Maybe. I think it's a fantastic illustration of anyone doing anything to provide something other than hype around the subject. An actual RCT? Too good to be true! The thirst is for fact vs. speculation, influencer blogposts and self-promotion.
That this RCT provided evidence opposing the hype, is, of course, irresistible.
I think that part is understated, while the part about the "steep learning curve" has been considerably overstated. Ive watched plenty of good developers with hours of cursor practice, using plenty of relevant context giving clear instructions wasting chunks of 5-10 minutes at a time because 4/5 times claude produced crap that needed to be undone.
That story doesnt sell though. Nobody will pay me consulting fees to say it. Hype has more $$$ attached.
Why should we be eager to find out that some new tech is going to undercut us and replace us, devaluing us even more than we already are?
Open source packages are the biggest productivity boost of my entire career, at no point did I think "wow, I wish these didn't exist, they're a threat to my livelihood".
Could it be there'd be less jobs because it would have prevented some businesses to bootstrap and stay afloat? That's also possible.
And this is really the thing with AI, some think it'll grow the pie, therefore increasing the demand for developers. Some think it'll just shrink the size of the piece of pie that's for developers.
If we could predict the former, and be sure of it, I doubt anyone would be against AI.
Aside that though, as a developer, you are also concerned with, is this actually making my job better/faster/more-fun. Or am I just forced into an annoying process of using a tool that doesn't help and will just cause more problems I'll have to deal with.
Prior to open source software development was awful. We spent all of our time re-inventing wheels - or, if we were lucky, licensing expensive badly shaped wheels from vendors and crossing our fingers that they would work (because we couldn't fix them if they didn't).
As a result, very few companies developed software - it's not worth hiring a software engineer if it takes them six months to deliver that initial login screen.
Open source started taking off about 25 years ago, and software productivity skyrocketed. The result was a multi-decade boom in jobs and salaries for us software engineers.
Still do.
At least back then it probably felt more like we were doing something new because 8000 versions of the same thing wasn't right at your fingertips.
I have no leg to stand on, I owe my career to proprietary software building on top of open source code, but lately I have been wondering what that alternate world would look like. As it stands, if you want to write software and sell it, you've got at most a couple of months before someone open sources a clone of your project. Sure, there's no guarantee it'll be anywhere near as good. Photoshop is still king. But one can't help but wonder what the proprietary IBM world would have been like.
Back then the answer was supposed to be OOP. With hindsight, the answer was open source - in particular robust library ecosystems like PyPI and NPM.
It looked more fun on paper, FWIW.
What if that happened across the entire industry overnight?
What if every single time you sat down to write a piece of software, someone immediately produced the Open Source version of it?
That's what people are hoping AI is going to do to us
Why would I support that
WordPress did that to the commercial CMS industry. Linux did that to the commercial Unix industry (RIP, Sun Microsystems).
> That's what people are hoping AI is going to do to us
I think those people are full of crap, and I mostly ignore them.
Open-source does replace entire companies, but typically with better products that other companies can build upon, and in the end, this expands the demand for software.
AI is expected to replace developers, and I believe that it will empower non-techies to develop highly-functional mockups, which will in turn increase the amount of (crappy) software on the market, which I believe will in turn increase the demand for people who can fix this software – if we multiply the number of apps by 1,000, even if only 1% of the apps need quality, this still means that we've increased the demand for software developers 10x.
So I actually think that, within a few years, AI will increase our job security as software developers. It's going to be a rocky few years, of course, but I think we'll be better off, at least until the rest of our dystopian timeline catches up.
Now, what I hate is the dishonesty of AI companies. I mean, as software developers, we strive on rigor: good documentation, compilers and JITs that speed up our code without changing its behavior, APIs and tools that work as claimed, etc. AI companies and their PR, on the other hand, peddle magic thinking and gaslighting. That's the exact opposite of rigor. And this keeps assaulting my sanity.
Should be looking for ways to work slower? I can go back to just one monitor.
If productivity gains mean that either me or a coworker is laid off, and my boss makes more money, why should I care about the that?
Having a job is not a favor that people with money are doing for employees, it is an exchange of services for money
I provide a service. So do most people. Anything that gives the people with money a way to avoid paying for my service is bad for most of us
"The tragedy is that although these groups of workers did suffer from the deployment of new technology, over time, the general population benefitted from cheaper clothes, lower prices and new jobs in other sectors of the economy."
Obviously, as a developer, I don't want to lose my career. But am I a hypocrite? Are you a hypocrite for buying anything that is now manufactured by automation, putting people out of jobs? I'd say yes.
It’s also probably not coming from a place of “I’m scared of AI so I want it to fail” but more like “my complex use case doesn’t work with AI and I’m really wondering why that is”.
There’s this desire it seems to think of people who aren’t on the hype train as “against” AI but people need to remember that these are most likely devs with a decade of experience who have been evaluating the usefulness of the tools they use for a long time.
The time is the early 2000s, and the Segway™ is being suggested as the archetype of almost all future personal transportation in cities and suburbs. I don't hate the product, there's neat technology there, they're fun to mess with, but... My bullshit sensor is still going off.
I become tired of being told that I'm just not using enough imagination, or that I would understand if only I was plugged into the correct social-groups of visionaries who've given arguments I already don't find compelling.
Then when somebody does a proper analysis of start/stop distance, road throughput, cargo capacity, etc, that's awesome! Finally, some glimmer of knowledge to push back the fog of speculation.
Sure, there's a nonzero amount of confirmation bias going on, but goshdangit at least I'm getting mine from studies with math, rather than the folks getting it from artistic renderings of self-balancing vehicles filling a street in the year 2025.
Yes, exactly!
I've spent way too much time trying to get anything remotely close to an LLM writing useful code. Yeah, I'm sure it can speed up writing code that I can write in my sleep, but I want it to write code I can learn from, and so far, my success rate is ~0 (although the documentation along the bogus code is sometimes a good starting point).
Having my timelines filled by people who basically claim that I'm just an idiot for failing to achieve that? Yeah, it's craze-inducing.
Every time I see research that appears to confirm the hype, I see a huge hole in the protocol.
Now finally, some research confirming my observations? It feels so good!
My team does two person PR reviews for example. We'd go a lot faster if we didn't or even just allowed a single reviewer. Similarly, we have no idea what the quality impact would be if we stopped, and what we gain by doing so. We are we not having a 3 reviewer rule for example, why not, two is an arbitrary number?
Unit tests... We'd surely go a lot faster if we didn't bother with them. Teams used to have some dedicated QA members and you'd rely entirely on manual testing. You can push a lot more code out. Was software in the 90s when unit tests and integ tests wasn't used buggier than today's software?
Now take AI, what is the impact of its use? It's not even obvious if it reduced the time it takes to launch a feature, my team isn't suddenly ahead of schedule on all our projects, even though we all use Agentic tools actively now. Ask any one of us and "I think it makes us faster" will be the answer. But ask us why we have a 2 person review rule and we'd similarly say: "I think it prevents bugs and improves the code quality".
The difference with AI now is that you pay for it, it's not free. Having unit tests or doing a 2 person review is just a process change. AI is something you pay for, so there's more desire to know for sure. And it also is something people would like to know if they can lower their headcount without impacting their competitive edge and ability to deliver fast and with good enough quality. Nobody wants to lower the headcount and find out the hard way.
Yep. It's been a "problem" for decades at this point. Business types constantly trying, and failing, to find some way to measure dev productivity, like they can with other types of office drone work.
We've been through Lines of Code, Function Points, various agile metrics, etc. None of these have given business types their holy grail of a perfectly objective measure of productivity. But no one wants to accept an answer of "You just can't effectively measure productivity in software development" because we now live in a data-driven business culture where every little thing must be measured and quantified.
What you’re seeing is a thirst for objective reporting. The average person only has the ability to provide anecdotes - many of which are in stark contrast to the narrative pushed by the billionaires pumping AI.
I don’t think anyone serious thinks AI isn’t useful in some capacity - but it’s more like a bloom filter than a new branch of mathematics. Magically powerful in specific use cases, but not a paradigm shift.
They're obviously talking about the METR paper, but the main takeaway according to the authors themselves was that self-reporting productivity increases is unreliable, not that you should cancel your subscription.
Nothing in that paper said that AI can't speed up software engineering.
Why are we responding to hype with nonsense?
I mean, the paper did provide tangible data that at least in their experiments, AI slowed down software engineering.
What they said is that it's not a proof that there's isn't a scenario or a mechanism where AI could result in speeding up software engineering. For that more research would be needed in measuring productivity of AI in more varied contexts.
For me at least, their experiment seem to describe the average developer's use of AI. So it's probably telling you that currently on average AI might be slowing things down.
Now the question is, can we find good data of outliers, and is it a simple matter of figuring out how to use it effectively, so we can upskill people and get the average to now be faster. Or will the outlier be conditioned on like, only for newbies, only for prototypes, only for the first X weeks on a greenfield code base, etc.
Edit: That said, the most fascinating data point of that study is how software engineers are not able to determine if AI makes them faster or slower, because they all thought they were 20% faster but were 19% slower in reality. So now you have to become really skeptical of anyone who claims they found a methodology or a workflow where their use of AI makes them faster. We need better measurement than just "I feel faster".
There is nothing in it for me, if I am more productive but earn the same and don't get any more time off
Why should I bother at that point?
1) If you are a salaried employee, if you are seen as less productive than your colleagues that use AI, at the very least you won't be valued as much. Either you will eventually earn less than your colleagues or be made redundant.
2) If you are a consultant, you'll be able to invoice more work in the same amount of time. Of course, so will your competitors, so that rates for a set amount work will probably decrease.
3) If you are an entrepreneur, you will be able to create a new product hiring less people (or on your own). Of course, so will your competitors, so that the expectations for viable MVPs will likley be raised.
In short, if AI coding assistants actually make a programmer more productive, you will likely have to learn to live with it in order to not be left behind.
That is to say: "Productivity" is notoriously extremely hard to measure with accuracy and reliability. Other factors such as different (and often terrible) productivity measures, nepotism/cronyism, communication skills, self-marketing skills, and what your manager had for breakfast on the day of performance review are guaranteed to skew the results, and highly likely, in what I would guess is the vast majority of cases, to make any productivity increases enabled by LLMs nearly impossible to detect on a larger scale.
Many people like to operate as if the workplace were a perfectly efficient market system, responding quickly and rationally to changes like productivity increases, but in fact, it's messy and confusing and often very slow. If an idealized system is like looking through a pane of perfectly smooth, clear glass, then the reality is, all too often, like looking through smudgy, warped, clouded bullseye glass into a room half-full of smoke.
Because productivity is hard to measure, if we just assume that using AI tools is more productive we're likely to be making stupid choices
And since I strongly think that AI coding is not making me personally more productive it puts me in a situation where I have to behave irrationally in order to show employers that I'm a good worker bee
I am increasingly feeling trapped between a losers choice. I take the mental anguish of using AI tools against my vetter judgment or I take the financial insecurity (and associated mental anguish) of just being unemployed
vs.
--- start quote ---
In a randomised controlled trial – the first of its kind – experienced computer programmers could use AI tools to help them write code.
--- end quote ---
Your quote is very representative of the magical wishful thinking most people have about AI: https://dmitriid.com/everything-around-llms-is-still-magical...
Gosh, I was conflicted, then you pulled out that sentence and I was convinced. :)
Alternatively: When faced with a contradiction, first, check your premises.
I don't want to belabor the point too much, there's little common ground if we're at all or nothing thinking - "the study proved AI is net-negative because of this pull quote" isn't discussion.
Your comment here is very representative of how quickly people who are AI skeptics will jump on anything that supports their skepticism.
we don't demand every developer pop Adderall though
Me: The person above literally pitches an unsupported belief against a study
You: it's pretty easy to believe your own experience over even a well-constructed "randomised controlled trial".
Really? Really?!!
As for "boosting your productivity", it's also what I'm talking about in the article I linked:
--- start quote ---
For every description of how LLMs work or don't work we know only some, but not all of the following:
- Do we know which projects people work on? No
- Do we know which codebases (greenfield, mature, proprietary etc.) people work on? No
- Do we know the level of expertise the people have? No. Is the expertise in the same domain, codebase, language that they apply LLMs to? We don't know.
- How much additional work did they have reviewing, fixing, deploying, finishing etc.? We don't know.
Even if you have one person describing all of the above, you will not be able to compare their experience to anyone else's because you have no idea what others answer for any of those bullet points.
--- end quote ---
So what happens when we actually control and measure those variables?
Wait, don't answer: "no, it's easier to believe yourself over a study".
See? Skeptics don't even have to "jump on anything that supports their skepticism." Even you supply them with material.
I've been banging this drum for over a year now: LLMs are deceptively difficult and uninituitive to use. Just one example: https://simonwillison.net/2025/Mar/11/using-llms-for-code/
What I'm willing to assert as fact, based not just on my own experiences (though they're a major role) but on observing this space for several years and talking to literally hundreds of people, is that LLMs can provide you a very real productivity boost in coding if you take the time to learn how to use them - or if you get lucky and chance upon the most productive patterns.
EDIT: I just saw you're the author of https://dmitriid.com/#everything-around-llms-is-still-magica... - that was a great piece! I think I may actually agree with you. I misinterpreted "magical thinking" as referring to something else.
Thank you!
> I think I may actually agree with you.
I was just going to write "see, you actually agree with me", but got hit by the reply rate limit :)
And I agree with >90% of what you write, so I was surprised that this bout took us to weird places.
Im sure they were completely genuine in how they felt, just as i am sure you are too.
In my youth, I would have argued this was bad. Now, I tend to agree. Not that studies are worthless; but they are just part of the accumulation of evidence, and when they contradict a clear result you are directly seeing, you need to weight the evidence appropriately.
(Obviously, replicated studies showing clear effects should be more heavily weighted.)
Everything is just shifting odds.
Either way, it's hard to interlocute if your misreading is an absolute conclusion that cannot be argued, then transmutated warranted skepticism.
Edit: SimonW? Really? I didn’t see the name but I didn’t expect you to be like that.
- Snark?
- Is "the issue" that anyone who claims any productivity gain is using magical thinking?
- How does the linked article "deal with" "the issue"?
- What title did they mention?
- What did they link to that has that title?
> Edit: SimonW? Really? I didn’t see the name but I didn’t expect you to be like that.
Like what? I think you're getting a bit emotional & personal here, I don't read anything remotely inappropriate into Simon's comment. Been here 15 years. OP's was odd for HN in that it admits 0 argument: if you think you have productivity gains, it's magical thinking.
My comment was in response to Simon’s reply to a user who posted an article. The title of the article they posted addresses magical thinking in AI.
Now whether that’s an opinion you share or not is not the point. Simon responded as if the user was only calling any perceived gains from AI as magical thinking which is not the case.
I’ll let you come back to that when you feel like it. Altogether, though it’s just disappointing to see someone who’s work I read often jumping to an emotional response when it’s not warranted.
I don't think the response from troupo that nayshins's personal experience is invalidated by a "randomised controlled trial" was well argued, so I imitated what I saw as their snarky wording with my own reworded version of it.
I do take the "AI isn't actually a productivity boost" thing a little bit personally these days, because the logical conclusion for that is that I've been deluding myself for the past two years and I'm effectively a victim of "magical thinking".
(That said, I did actually go to delete my comment shortly after posting it because I didn't think it added anything to the conversation, but it had already drawn a reply so I left it there.)
Working in security I often feel the same way and let’s be fair in the grand scheme of things it’s not that big of a deal.
You may just as well have. I, for one, am absolutely ready to re-evaluate any and all approaches I have with AI to see if I am actually more productive or not.
But moreover, your own singular experience with your own code and projects may make you more productive. We don't know if it does because we don't have a baseline against which to measure.
But even moreover over that moreover is that we don't even have a question "does a single senior engineer's experience with his own code and approaches can be generalised over the entire population of programmers?" Skeptics say: no (and now have some proof of that). Optimists loudly say: yes, of course, and dismiss everyone who dares contradict out of hand.
the psychological effect reminds me a bit of slot machines, which provide you with enough intermittent wins to make you feel like you're winning while youre lose.
I think this might be linked to that study that found experienced oss devs who thought they were faster when they were in actual fact 20% slower.
At the game of producing garbage slop? Probably yeah.
This sounds like in there will be a race between this kind of booby trap tests and AIs learning them.
In quite a few interviews in the last year I have come away convinced that they would have performed far better if they had relied on their own knowledge/experience exclusively. Fumbling with windows/tabs, not quite reading what they are copying, if I ask why they chose something, some of them would fold immediately and opt for something way better or more sensible, implying they would have known what to do had they bothered to actually think for a moment.
I put down "no hire" for all of them of course.
What's the goal of this? What are you looking for?
In the real world, you hit problems that the LLM doesn't know what to do with. When that happens, are you stuck, or can you write the code?
As it happens, this meant when candidates started throwing AI at the task, instead of performing that magic it usually can when you make it build a todo app or solve some done-to-death irrelevant leetcode problem it flailed and left the candidate feeling embarrassed.
I really hope AI signals the death knell of fucking stupid interview problems like leetcode. Alas many companies are instead knee jerking and "banning" AI from interview use instead (even claude, hilariously).
How exactly did you outperform? Show, don't talk.
How is anyone supposed to understand what this means?
Given the ambiguity in your description and lack of actual code it’s hard to take you seriously.
But then when I really think about it usually they're just bullshitting out being purposefully vague, using terms that don't mean anything precise in order to avoid actual criticism.
Doing bad things faster might feel more productive to you, but it doesn’t mean that you are delivering more value. You might be, but the metrics you have shared to not prove that.
There are many things in this world that could be fairly described as "more productive" or "faster" than the norm, yet few people would argue that it makes those things a net benefit. You can lie and cheat your way to success, and that tends to be successful too. There are good reasons society frowns on this.
To me, focusing only on "I'm more productive" while ignoring the systemic and societal factors impacted by that "productivity" is completely missing the forest for the trees.
The fact that you further feel that there isn't even a point in engaging on the topic is disturbing considering those ignored factors.
actually based on your own admission this is not what you're doing...
People who boast about AI enhanced productivity seem to always forget to mention.
I’ve watched coding change from Cursor-esque IDEs to terminal based agentic tools within months.
I still suspect the vast silent majority of professional software devs haven’t integrated any, even Cursor-style, AI tools in to their main gig.
And I reckon that’s completely rational, for those that have made this choice explicitly.
Early adopters of AI tools are making a speculative bet, but so far most of them seem happy with the return.
This piece, however, only focuses on time spent on a task that could be done both ways. Even there, it falls short. Let's assume this study is correct and a specific coding task does take me 19% more time with AI. I can still be more productive because the AI doing some of the work allows me to do other tasks during that time.
I do worry about atrophy of my mind outsourcing too many tasks, admittedly. But that's a different issue.
It's similar to the story of the development of vehicles and how even though we move much faster we spend a greater amount of time in transit. My mom used to lament how annoying it was to have to drive to the grocery store because when she was younger and not everyone had cars the store came to you. Twice a day, in the morning and the evening the "rolling store" would drive through the neighborhood and if they didn't have what you needed right then, they would bring it on the next trip. We are finally coming back full circle with things like Instacart but it's taken a solid ~60 years of largely wasted inefficient travel times.
I think you're onto something with this take, based on my own experience. I definitely agree that my RPE seems lower when I'm using AI for things, whether it actually is making me more productive or not over the long term remains to be seen but things do certainly "feel" easier/less cognitively demanding. Which, tbh, is still a benefit even if it doesn't result in large gains in output. Putting in less cognitive load at work just conserves my energy for things that matter - everything else outside of $dayjob.
Also I do not think AI is falling apart and I do not think the productivity loss is surprising. It's something I've noticed a lot, you have to check AI output line by line thereby making you slower. AI makes tasks easier, not necessarily faster. And because people are lazy and love making things easier it's still going to be adopted no matter what anyone says.
They forgot about training and practice. Try it with airplanes. "The results surprised us. Given Boeing 777 they thought they will get from London to New York in 7 hours, but they actually couldn't even get off the ground. In fact they couldn't even start the engines."
If your "AI Assistant" is as hard to operate as a trans-Atlantic flight on a Boeing 777, perhaps it's not a very good assistant?
Since then, using Gen AI tools to learn + write code, I have deployed a functioning full stack application to the web. NextJS on Vercel, with a backend server also deployed running on Python, and a Supabase DB. Is it the best application ever making loads of money? Definitely not. Are there things wrong with it? Absolutely (although I promise I'm not exposing sensitive env vars and API keys to the web). Did the first versions look like absolute ass as I clumsily figured things out and made bad mistakes? You bet. But it's a functioning app that does some useful things and has real users.
I would never have imagined doing this in a million years prior to Gen AI.
Do some devs see mixed results depending on how they're using the tools? I'm sure. Is Gen AI overhyped broadly speaking? Probably so. But when I see people say it's a delusion, waste of resources, and everyone is wasting their time on it ... for me, it just doesn't line up.
Best hope ChatGPT plagiarized that code from people who properly sanitized user input, otherwise it might be vulnerable to SQL injection, XSS, etc. If such holes exist, it may be tough to resolve them without a professional.
> But when I see people say it's a delusion, waste of resources, and everyone is wasting their time on it ... for me, it just doesn't line up.
I don't think anyone would say LLMs are useless, but there are people out there comparing the birth of LLMs to the Industrial Revolution, and the advent of the Internet. Such claims are preposterous.
I think the widespread use of LLMs will ultimately be undone by a few factors:
1) Due to the way LLMs function, they will never be totally cured of 'hallucinations'. They are accurate a lot of the time, and apparently useful in those times. But you can never really trust its output. That is an exhausting problem, and will burn people out.
2) Prices for LLMs are artificially very low right now, these companies are burning investor money to keep going. The 'cost' is going to have to go way up for these LLM providers to be sustainable. I put cost in quotes because it will be some cost in subscription prices, but also cost in terms of ads that will be inserted everywhere, and other degrading forms of monetization.
3) Models are running out of good training data. They'll either hit a wall, or continue training on content that was itself LLM generated, which will go the way of Habsburg Jaw in a jiffy. Ensloppification is real.
I'm sure there will be niches for LLMs in the future, for those who can afford their true cost. But this cram-it-in-everything, available-everywhere frenzy will probably be looked back upon with deserved embarrassment.
What an utterly bizarre question. Yes, by definition being more productive means doing more work.
I've seen this firsthand. It's so easy to produce "content" now, like emails, presentations, etc., that more of my time now is just sifting through AI slop looking for signal.
I'm on a board for a nonprofit, and we've seen a person that is upset with the organization send dozens of AI-generated quasi-legal demands full of hallucinations of rules and laws that don't actually exist. So now we're forced to pay a real human lawyer to deal with a denial of service attack of ChatGPT slop and are spending lots of time just dealing with the onslaught. Our productivity has crashed as a board due to a single bad actor enabled by ChatGPT.
At work, I'm also getting buried in AI sales calls and AI resumes. It's getting more and more difficult to find signal in the noise.
It wouldn't surprise me if there's a lot of people who are experiencing lower productivity for similar reasons.
10% would mean a 10 trillion in productivity, 1% is 1 trillion.
It implies that even if some tasks are better with AI that it might not be so simple for us to judge which method is better for which task.
And even then, the web was over-hyped at various points in its life. And it took a terrible crash to sort out what was real from what wasn’t. It’s not that the web wasn’t absolutely transformative (it was!) but it wasn’t what a huge number of proponents thought it would be either.
You’ve seen the posts and articles: “I’m inherently better than AI, coding faster and smarter.” Sure, maybe for now—especially if you’re half-assing your AI game. But that edge? It’s vanishing fast.
Stop being so self indulgent.
Become proficient with the tools available to you or get left behind.