My take is the product has been very useful for coding (PMF) for months. But it’s certainly not useful at any cost…
My take is the product has been very useful for coding (PMF) for months. But it’s certainly not useful at any cost…
And that's just one inflection point. We've had several and there are many more on the horizon. So while I could be convinced that ROI is maybe not even positive today despite the ridiculous enterprise spend, it's perfectly rational to pave the way today for what's coming over the next few months let alone years down the line.
Somewhat oversimplifying; writing software and building apps was a bottleneck - now it is not. What is the next bottleneck that LLMs can solve? Is there one? And is there enough publicly available data to solve it repeatably at scale? Or did we just automate stack overflow searches and now we’re stuck again?
Or is the endgame of this innovation cycle the complete removal of interaction with machines through code? Will we simply interact with machine coworkers purely through natural language? Can an LLM make PowerPoint slides and run a meeting? So far not seeing much progress on that.
Nor is it a good reason to think there will be more.
https://substackcdn.com/image/fetch/$s_!_ZW2!,f_auto,q_auto:...
If you can talk about my irrational skepticism (because I said that "we don't know the future", I suppose?), can I talk about your total lack of common sense?
Because the economy has been growing in the last decades does not mean that it will keep growing for the next decades. Because LLMs have been improving in the last few years does not mean that they will keep improving in the next few years. Maybe, maybe not, your guess is as good as mine. If you know the future, put your money where you mouth is and invest everything you own in LLM companies.
Your overwhelming evidence is about the past: it has been improving in the past.
In any case, maybe I was too subtle. I was talking about Mythos, a model that continues the trend, but which is not available to the public yet. The "overwhelming evidence" is the testimony of the people who have used it. The irrational skepticism was people who don't believe that testimony. In other words, we do know the future, because we know that model and others like it will come out soon.
I just have an issue with all the people saying "I predicted this 10 years ago" (implying something like "you should listen to me, I make good predictions") while conveniently forgetting all the things they predicted wrongly, or the survivorship bias.
We don't know that AIs will continue improving at the pace they have, because we don't know the future. Some people will guess right, some won't. And those who guess right will be tempted to believe that they guessed right because they are more clever. All we can say is that it is possible that it improves, and it is possible that is stops improving.
Sure, it might start to slow down, but even then we will likely see a doubling in the next 10-15 years.
https://substackcdn.com/image/fetch/$s_!_ZW2!,f_auto,q_auto:...
In other words, most of the prompting will also go away.
Ultimately software is everything these days and the economics make the demand insatiable. We've gone through many cycles of "X" but on computers/web/mobile. There's going to be a massive amount of "X" but with AI companies that will need engineers.
Or at least this is what I tell myself to sleep at night.
Ultimately, we'll need UBI or large scale cuts in working hours or similar if AI progresses to the point of mass unemployment - the alternative would be massive social unrest. In the meantime I expect to keep doing better than average.
I think it was clearly useful for months to people who had tried it and taken the time to understand it, but now that knowledge has spread to the point where wallet holders are convinced it's not just passing fad or hype so now pmf can be "claimed".
I agree it's weird to say "those people have pmf" though, usually it's something you define for yourself
I'm not sure if this runs counter to your point or not, but: I don't see any future where LLMs aren't a core part of Software Engineering. The horse is out of the barn. There is no going back.
And I don’t even necessarily disagree with OP! It’s more like the competition is shifting so quickly that your competitors could undercut your PMF in a blink of an eye.
But my guess is that the cost of SWEs themselves mean that the more expensive ones will be worth the delta to most companies.
But time will tell.
people -> programmers, I haven’t met a non-developer who reports getting more time out of current AI platforms than they put in. If anything I’ve anecdotally heard the opposite, introducing AI at work creates so much slop (output) it takes more time to process it all without a tangible bump in overall productivity
Thats why most here shouldn’t engage in the discussion - they parrot on about benefits without identifying and articulating the costs and moreover how it affects the firms financial position.
"I’ve called November 2025 the November inflection point because that was when GPT-5.1 and Opus 4.5, combined with their respective coding agent harnesses, got good—good enough that we’ve spent the last six months adapting to agent systems that can reliably get useful work done."
Not saying this trend will do the same, just that the industry adopting something doesn't guarantee its success.
By comparison almost all tech companies I know have leaned heavily into AI.
If I make an argument and you disagree that's fine with me, provided I didn't use misinformation or sloppy thinking in making that argument.
My root comment simply represented my two cents about the current post. I don't think anything about the post is outrageously incorrect or anything, just somewhat confusing. You're a very prolific contributor in this community and I don't think me or anyone else that welcomes your takes expects everything you write to rock our collective socks every single time, anyway.
52 on AI misuse: https://simonwillison.net/tags/ai-misuse/
149 on the unsolved challenge of prompt injection: https://simonwillison.net/tags/prompt-injection/
40 on slop: https://simonwillison.net/tags/slop/
If you want an "LLM evangelism blog that rarely, if ever, has any critical analysis that isn’t pro-industry" there are plenty out there. I'm not one of them.
Many people still think AI coding agents are slop on steroids despite all the current hype around AI actually shipping functional products.
(And that's after taking into account the METR paper that says engineers over-estimate their productivity with these tools.)
I have plenty of doubts about AI delivering on its promises outside of coding. I don't write about AGI because I think it's science-fiction hysteria. I write about slop precisely because it represents a mis-use of AI that demonstrates people completely misunderstanding what it's useful for.
"Many people still think AI coding agents are slop on steroids despite all the current hype around AI actually shipping functional products."
Oh yes, tons and tons, especially on HN. But the plural of anecdote is not data. Enterprise spend speaks for itself. You are using AI-coded functional products all the time. Do you want like a diff history for the Google codebase or something?
> I’ve called November 2025 the November inflection point because that was when GPT-5.1 and Opus 4.5, combined with their respective coding agent harnesses, got good—good enough that we’ve spent the last six months adapting to agent systems that can reliably get useful work done.
Claiming a grand inflection point based on your own personal usage is very anecdotal.
Compelling anecdotes are not even the main source of evidence. Look at the enormous body of work on measurement of these systems. I always point people to epoch capability index as a good summary statistic of capabilities or METRs time horizon data which has now been topped out. They had a recent updated to the dataset, after which the corrected plots pointed to an even faster acceleration than before.
That's exactly what I'd expect people who are driven by hype and FOMO and YOLO and anecdotal evidence to do.
> resulting in systemic collapse.
Many people are noting the system is collapsing. Maybe it's not going as quickly as you expect, but there's definitely evidence of this from increased service outage frequency, billion dollar notes being passed in a circle between companies, open projects refusing AI contributions entirely because they're overwhelmed by crap, Sam Altman begging governments to force citizens to buy their product through "universal basic compute", etc.
> Look at the enormous body of work on measurement of these systems.
It's certainly possible to measure anything. Benchmarks are a form of evidence but they famously a) don't represent reality and b) can be easily gamed.
Not at this scale.
> Many people are noting the system is collapsing.
On HN? any piece of evidence to support this? service outage frequency is not a sign of systemic collapse. billion dollar notes passed in a circle is brought up a lot and misunderstands how finance works. "open projects refusing AI contributions entirely because they're overwhelmed by crap" is not a systemic collapse, its not being able to adapt to a new world with new challenges. Btw "slop" is getting less and less sloppy.
> Sam Altman begging governments to force citizens to buy their product through "universal basic compute", etc.
Very interested in some citation detail that sounds like a headline quote of something more complex.
> It's certainly possible to measure anything. Benchmarks are a form of evidence but they famously a) don't represent reality and b) can be easily gamed.
I mean I work on benchmarks for a living I can tell you both of these things are true but only partially, and in aggregate they all tell a consistent story. Not to mention, static OSS benchmarks are not what these companies rely on. They have live traffic, ability to run A/B tests, full conversation traces, to ignore this is pretty incredible.
If you want indisputable, data-driven information about the state of the LLM world I guess you can wait for a peer-reviewed academic paper?
I did try to provide credible links to back up my assertions about things like enterprise pricing changes.