Time will tell of course, and it’s early, but inflection points do exist with progress.
Time will tell of course, and it’s early, but inflection points do exist with progress.
“But it’s different this time” - several people, several times over the last couple of years.
This is not at all a dig at you, I’m very sorry if it reads that way. My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change. Just yesterday I had gpt 5.3 generate completely awful code for the Cinema 4D Python API. Also an anecdote. But for all of the people saying they are truly intelligent and truly reason, they still make obvious mistakes, write around problems, fail entirely at architectural decisions, fail at random, generate FAR too much code.
And no amount of harnesses, methodologies, loops make much of a difference. If you listen to people on the internet they say it’s all working. You listen to people on the job and they mostly say it’s creating tech debt and a review bottleneck. Also burnout, so much burnout.
I think LLMs are mediocre. I think it’s fine they’re mediocre. You can work with low expectations. But the hype cycles are so tiresome.
Then the incremental improvements did, in my experience, cross some kind of threshold in late 2025 where the things became useful. It is of course anecdotal and personal judgment. But I asked LLMs to implement a small feature in my codebase (my usual test) and finally it produced code I was happy with. They've also been able to locate and diagnose a problem based on logs. In my view it's now a markedly different level of capability than we had a year ago, though I would call the previous two years equally useless.
Yes, as a product gradually improves there will always be many people for whom version X didn't work well for them and version X+1 does. It turns out that Opus 4.5 and GPT 5.1 were larger than average improvements that cross that threshold for a significant number of people.
My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change.
If your claim is that there's no substantive difference between Sonnet 3.5 and Fable, then we live in very different worlds.
Even 3.7. I remember when it came out and people were claiming that that was now the model that was going to replace engineers. Cue Fable years later and people still claim that this one is the one.
How does GPT-5.6 Sol or Claude Fable 5 or Claude Opus 5 do on that Cinema 4D code?
Why would you attempt to use GPT 5.3 to generate code today and form an opinion on that basis?
I do not think it is even still available in Codex, I believe it only has the smaller, distilled GPT 5.3 Codex Spark.
What shape did that evidence take?
The key thing is that it's easy to contrast the old way vs the new way and the evidence become obvious.
The thing with LLM tooling is that they're not reliable. I can do fine with risks, but only when there's a way to manage it so that if the disaster happens, it's practically a black swan event.
Typing more code or solving one task has never been the core problem. The core problem has always been to encode a whole system into the computer AND then provide a control interface for it. It requires both an understanding of the system you want to encode (especially how it behaves over time) and empathy to know what would be the best control interface for the users.
That understanding does not rely on the amount of code, and the best control is found through communication.
If we take the following project that you did:
https://simonwillison.net/2025/Jul/17/vibe-scraping/
An understanding of the system could be the following: A conference schedule consisting of events (time, place, speaker, description,...) stored or presented in some format. The interface would be: A web app with a mobile first UI that presents the information in an accessible manner (highlighting, filtering, exports,...).
A relatively quick (I haven't tested it), would have been to open the web inspector and extract the data using the dom API (requires knowledge of the dom api and a desktop browser), put the data into some json or a tsv file, then write a php script or a python script and then serve that. The interface could have been built with the standard elements of some css framework (bulma?).
Not saying the above is better. But the thing is that is doable from even a raspberry pi. And more it's repeatable and extensible. And the individual piece of knowledge are reusable in different situation.
Pro: "The evidence hasn't shown up in statistics yet; it's too new!"
Con: "And won't this destroy maintainability?"
Pro: "Show me the maintainability disasters caused by AI."
Con: "I can't yet; it's too new!"
Both sides are playing the "it's too new" card when asked for actual evidence to prove their claims. In fairness, it actually is too new for there to be much statistically-valid data, especially if the inflection point was November 2025. So both sides are trumpeting their position, neither with actual trustworthy data.
Everybody has their anecdote. Nobody has data yet.
AnimalMuppet in their reply to my OP comment makes the point that these things haven't been around long enough to measure end-to-end productivity gains and come to a conclusion either way, and I agree with that. But we can still at least be measuring something.
For example, in my case AI has allowed me to write 10x more LOC than I usually would in a similar amount of time. But having to review it all, I've also deployed 1/4 the number of releases I normally would in the same period. By one measure I'm more productive, by another I'm less productive.
People could claim to be more productive by skipping the review. But in that case did AI make you more productive or did you lower standards? People could say they're using AI to do the review but is the impact of that being measured and has that caused more or fewer bugs? If more bugs, has the time to fix those been factored into overall productivity? In my experience it's common for people to eagerly count immediate productivity gains and discount long-term productivity sinks.
For this reason I think case studies are the best convincing thing, because they properly contextualize the usage and consider a longer-term window. They're also backwards looking instead of in-the-moment, so have the benefit of hindsight. But they're harder to come by and we probably won't see any meaningful case studies for a thing that people say happened in November.
But the very least people can be doing is just defining what they mean when they say "productivity" because otherwise everyone is talking past one another.
Very smart people aren't immune to being worn down over time
Meanwhile the guy who leaned in a year ago and gave up reading the output is beginning to see work grind to a halt and throwing more agents at it is increasingly not working.
You can see these tropes all over social media near constantly.
You should stop using social media as your yardstick.
Yes, don't believe people posting on HN.
But in all seriousness - you can even bring up Pope himself. Don't care. Show me data, show me the leaps our software made with all this 100x productivity boost. Show me a myriad of better LLVM projects, new usable kernels, and so on. Show me sharp decline in bugs and defects in existing projects.
Keep blog posts and HN comments.
it's useful for scaffolding but after that I'm not sure how you could rely on it without being in the loop and directing how the code should be like
Jason Turner gave an excellent talk at last year's CppCon explain how he thinks tools can be used to make generative AI coding assistance safer and more productive. https://www.youtube.com/watch?v=xCuRUjxT5L8
Now it is.
Of course it doesn't seem that way to you. Preachers view themselves as spreading the good word, they don't see how annoying it is being preached at