2026? 5? 4? 3?
Heard this one way too many times.
2026? 5? 4? 3?
Heard this one way too many times.
And myself I keep making comparisons between AI and the progress in 90s video games where every minor improvement got called "photo realistic" and then forgotten with the next game engine: https://archive.org/details/nextgen-issue-26
So I'm not gonna say "this is it" when the software quality really matters, and I absolutely won't speak to progress (or lack of it) outside of software.
But I will say "you can look around and easily see small businesses using AI to generate posters, quite a lot of small business software and websites are in the same category: the mistakes are real but increasingly don't matter".
I think it would start to matter once again. People will get fed up of AI posters and art. I think they already are...and once some threshold is crossed, the business won't dare to use AI generated assets/designs.
Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.
Agreed, but will this look like a meme/fashion cycle? If so, re-prompt each year with a different look. Yes, there are still issues here, a friend found an image he was amazed was AI generated, but to me it was obviously so, so I showed him a screenshot of ChatGPT making something just it and included my prompt:
create image: hand drawing of cute springer spaniel puppy looking sideways, various geometric shapes drawn in layer behind and in front of the puppy, all done in style of 7 year old using crayons with mediocre colouring-in skills
As I said to them: yeah, the line thickness feels AI, to me, the bad colouring-in scribbles feel like just the art style it was propmpted with
it's like: it gets the big picture of the composition, and it knows how to colour in badly, but it doesn't know how to draw a dog as badly as the colouring in
> Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.We're better at recognising patterns full stop. All biological brains are, and needed to be better than the current state of the art in machine learning because if a living organism was as poor at learning patterns as the SotA in machine learning, the organism would starve to death before being able to pick up anything and eat it.
AI also has a second disadvantage, because there are so few models: the laziest of ChatGPT "thinkpiece" blog posts being everywhere is hard to miss, and 5000 fake bloggers all prompting the same model with "find biggest news story of today and write a blog post about it in a way that maximises my ad revenue" will get 5000 almost identical posts. This will remain true while each instance of the most commonly used AI fail to talk to each other in a way that at least mimics them collectively getting bored with writing the same thing 5000 times, it does not depend on e.g. quality.
Wrong analogy. A marketing item wants you to look at it. It takes advantage of an involuntary impulse. A repetitive boring AI artwork will not trigger the impulse to look it. So it is not about "caring"....
Right now GenAI content a bad thing, cringe, a sign of low-value or lack of attention to detail. I don't want to play even a free game if I think it was made by someone else prompting an AI, and that's despite liking the output when I do the same for myself.
If AI output is normalised and becomes simply "boring" or "mundane", those negatives must have also gone away.
After that (assuming there is an "after that", I don't want to bet either way), people would, as per your argument, need to un-boring them.
I have bad news for you if you think the video game industry isn't widely embracing generative AI. Companies try not to say it because of the current backlash involved but I truly believe Tim Sweeney isn't wrong[1]. I think in 5 years it will be a quaint idea to be AI-vegan and it'll be similar to how people resisted smartphones (I was one of them) until they became inevitable.
[1]https://www.techspot.com/news/110410-epic-tim-sweeney-ai-lab...
5 years is an eternity with the current rate of change of AI. Its development is already turning into an international geopolitical issue, even though the implications for mere task-level economics have yet to settle.
Even if it wasn't, 2031 (ish) happens to be roughly when a lot of long-term exponential growth trends all happen to reach points that suggest the assumptions behind them have to break, like more than 100% of electricity being made by PV etc.; the only thing I am confident of about AI is that even its current trend line on METR were to continue it becomes physically unmeasurable sooner than that. This also means nobody will pay Tim Sweeney so much as a single Zimbabwe Dollar in royalties for games written in the Unreal Engine.
Half the game logic I played on Kongregate back in the heyday of Flash games is now one-to-few shot prompting on Claude even if you don't pay for it, and even if you don't pay for ChatGPT you can get assets at that scale pretty quickly. I know 'cause I've tried. It's just, like every blog post whose headings contain emoji, every time there's a 6-7 word paragraph a little too bombastic about how important something is, every time I read or hear "you're absolutely right" or "delve" or "nuance", and yes hear even when the lips those words pass through happen to be human…
…the way most people prompt these models, I can tell. Most people don't bother to ask it to speak like a posh victorian, or draw poster art like it's the 60s, and they don't check and reject when the shadow on the foreground of the moon base doesn't match the Earth in the sky. And because humans are lazy and greedy, so even if the visual models gets as good as the finest artists and the language models as good as Nobel laureates, we're going to learn the style of whatever the default output of those models is when prompted by lazy and greedy humans. And we'll hate it, because it's a cliché, and we avoid those like the plague.
Handling highway driving with lane changes was great when it got there years ago, but just in the last year or so it has gone from a nice to have to “from now on I will never buy a car that can’t do this”.
AI has hit some milestones for replacing work as well. There’s still many more to go and maybe some of them will never get hit (much like I don’t think a coast to coast drive with zero interventions during winter conditions is ever going to happen) but there are points at which it forever meaningfully changes some field of work. I think it’s there for writing code.
Half-and-half. I'm not denying that self driving cars (and LLMs) are improving, I'm comparing it against the standards set by the biggest proponents. But yes, I have heard basically the same thing you just wrote for the previous several major releases of FSD.
Where we agree is that, while you are a fan, you do explicitly give as an example of something you think it will never do, something very close to what Musk has promised:
"Ultimately you'll be able to summon your car anywhere … your car can get to you. I think that within two years, you'll be able to summon your car from across the country. It will meet you wherever your phone is … and it will just automatically charge itself along the entire journey."
- Musk, in Jan 2016: https://en.wikipedia.org/wiki/List_of_predictions_for_autono...(That said, I think Tesla's FSD will never get there, not that it's impossible. The way Musk is behaving, there's going to be a financial scheme named after him in whatever passes for a textbook in 20 years, and it won't be the positive kind of example).
I'm mentally prepared for the next US administration to exact retribution on his companies, and I expect FSD will be neutered after that. Hopefully other car companies are able to catch up. I'm more of a self-driving fan than a Tesla fan, so as long as the thing works as well as what I have now, I'll be fine with it.
Yes. He was before and remains so to this day, but he did so then, too.
> I need it to safely drive my family on my daily errands or weekend trips, which I now consider solved.
This is probably unwise. As with LLMs, the statistics suggest a spikiness in the intelligence, with it being mostly good but also sometimes still making some very odd mistakes that humans would essentially never make.
As with your other comment, you know that if it gets into a crash you're responsible; while you consider this a win, I suggest waiting until the company you buy the car from (in this case Tesla) puts their money where their mouth is on quality and takes liability for crashes due to the AI upon themselves.
(The Cybercab was supposed to be sans-steering-wheel and sans-pedals, which would be a sign of that level of confidence, but the ones spotted in the wild at least sometimes seem to come with the wheel, which suggests they're still not there yet: https://www.vehiclesuggest.com/cybercab-with-steering-wheel-... https://www.carsguide.com.au/car-news/real-tesla-cybercab-sp...)
> Hopefully other car companies are able to catch up.
From the stats I've seen, they're much closer to the goal, relatively smoother/less spiky all-round driving intelligence. They may not be as impressive at their best, but when they fail the failure modes are themselves much safer.
The usual HN nits at this point are that the data is unreliable/biased and that the way I’m using it isn’t supervised enough. I’ll admit to the latter. In my experience the mistakes tend to be navigational but I’ve seen a few of the “nearly drove through the lane closed gate” videos so I definitely keep a closer watch when there is complicated highway stuff going on. I also have my foot at the ready for when I go through the gate to my community, similar failure mode and about 1% of the time it forgets to wait for the gate to close and reopen.
I look forward to the other car companies catching up, competition and more options are good. Elon is a loose cannon, so I need alternatives if Tesla ceases to be.
Tesla have been claiming the stats show superiority of their AI vs. humans since at least 2016: https://techcrunch.com/2016/07/06/tesla-says-drivers-using-a...
The lesson from the Datasaurus should be to ask detailed questions, e.g.:
• exposure bias: is the AI used more or less in certain conditions, on certain kinds of roads
• sharpshooter: shouldn't need to ask this one of a public company, but given regular headlines: how many crashes happen a few seconds after it switches itself off?
• driver demographics: most severe accidents by humans are due to impairment or being a new driver, is the AI actually better than an experienced-and-not-impaired driver?
• the nature of the failure modes: as a cyclist I've been hit by a car that stopped at a junction and pulled out into me, this was not fun but no serious injury; this is very different to any of the highway speed fatalities.
> Separately from the data, it is just so nice to hit a button and sit back and relax for the entirety of the drive.
No doubt. Myself, this is why I have been happy to live in a city with a good public transport network and use that instead of driving.
> I look forward to the other car companies catching up
My point is that they are in some important senses already ahead. Less optimised for the headline, more optimised for the real dream.
Usable public transport would be nice. It was unusable for various reasons where I used to live (Seattle/Bellevue) and is pretty much non-existent where I currently live (Socal / Orange County), but everything else about the location is perfect for me, so I'm willing to give that up.
> they are in some important senses already ahead. Less optimised for the headline, more optimised for the real dream.
Nice! I look forward to trying them out when they are available and working here. Rivian claims they'll achieve it next year, and one of their form factors (R1S) is perfect for my family so they'll be my next option if things go downhill with my Teslas.
It's not FSD until the human is no longer responsible.
This half measure bullshit is a joke.
It's just absolutely crazy to me that you trust this experimental feature more than the manufacturer does.
And this is not anecdotal, there are enough reports that an investigation is ongoing: https://autos.yahoo.com/policy-and-environment/articles/tesl...
I keep saying, FSD being marketed as FSD is going to get people killed and I can't believe more is not being done to prevent this.
I just spoke to a fried who is a headhunter and who's been trying to automate his processes for a while (he likes to fiddle and certainly has skills, but he's not an engineer). He kept trying, but it just wasn't good enough.
Now he said with GPT Work and Sol, it worked, but the key point is: all of it suddenly worked.
The problem was one of reliability, of handling edge cases. All previous attempts / model-harness-combinations were too brittle and needed too much observation and fiddling - cheaper to do it yourself.
Now he says "I don't know why I would ever hire a recruiter [the folks doing the cold outreach] again. I can focus on the candidate screening and acquiring projects, everything else is fully automated".
This doesn't come from an engineer or an AI lab, but a technically inclined power user, and I think this is where things get interesting.
This is like microprocessors in the 80s. Sure they double in capability every 18 months but the start is so pathetic it will be 30 years before they are good enough for everyday tasks.
It's cool that 'regular' people can now create solutions to many small problems, and automate stuff - genuinely a step forward. Like Excel, only vastly better. But for bigger projects, real software engineers know that what LLMs do today is only a tiny part of development. And it solves it in a way that might well make the rest of the lifecycle a lot harder. It's like that saying about tools that make easy things easier and hard things impossible.
Currently in a re-org justified by AI, AI changing the roles people will need to play. It’s disturbing how much content in the materials about the new org structure and roles and whatnot is clearly ai generated and contradictory. We’re laying off about half of 500 people.
It gets better: The internal AI gateway chat thingy where you can ask questions has AI autocomplete that pops up after your write more than 10 characters. What the fuck!?
would a great candidate get excited about an AI agent reaching out to them? or would it be the desperate or clueless one?
For all other roles the first few messages are similar: this is <role X> with key challenges a,b,c. You seem to be a good fit because of d,e,f. Would you be open to explore this? This requires relocating to <place>.
Traditionally, this was done by entry-level people. In either case, this isn't the person who will jump on a call with you.
"Desperate" and "clueless" seem very strong words here in reference to someone who gets actively approached from a recruiter.
Time will tell of course, and it’s early, but inflection points do exist with progress.
Very smart people aren't immune to being worn down over time
Now it is.
it's useful for scaffolding but after that I'm not sure how you could rely on it without being in the loop and directing how the code should be like
Jason Turner gave an excellent talk at last year's CppCon explain how he thinks tools can be used to make generative AI coding assistance safer and more productive. https://www.youtube.com/watch?v=xCuRUjxT5L8
Meanwhile the guy who leaned in a year ago and gave up reading the output is beginning to see work grind to a halt and throwing more agents at it is increasingly not working.
You can see these tropes all over social media near constantly.
You should stop using social media as your yardstick.
Yes, don't believe people posting on HN.
But in all seriousness - you can even bring up Pope himself. Don't care. Show me data, show me the leaps our software made with all this 100x productivity boost. Show me a myriad of better LLVM projects, new usable kernels, and so on. Show me sharp decline in bugs and defects in existing projects.
Keep blog posts and HN comments.
Of course it doesn't seem that way to you. Preachers view themselves as spreading the good word, they don't see how annoying it is being preached at
“But it’s different this time” - several people, several times over the last couple of years.
This is not at all a dig at you, I’m very sorry if it reads that way. My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change. Just yesterday I had gpt 5.3 generate completely awful code for the Cinema 4D Python API. Also an anecdote. But for all of the people saying they are truly intelligent and truly reason, they still make obvious mistakes, write around problems, fail entirely at architectural decisions, fail at random, generate FAR too much code.
And no amount of harnesses, methodologies, loops make much of a difference. If you listen to people on the internet they say it’s all working. You listen to people on the job and they mostly say it’s creating tech debt and a review bottleneck. Also burnout, so much burnout.
I think LLMs are mediocre. I think it’s fine they’re mediocre. You can work with low expectations. But the hype cycles are so tiresome.
Why would you attempt to use GPT 5.3 to generate code today and form an opinion on that basis?
I do not think it is even still available in Codex, I believe it only has the smaller, distilled GPT 5.3 Codex Spark.
Then the incremental improvements did, in my experience, cross some kind of threshold in late 2025 where the things became useful. It is of course anecdotal and personal judgment. But I asked LLMs to implement a small feature in my codebase (my usual test) and finally it produced code I was happy with. They've also been able to locate and diagnose a problem based on logs. In my view it's now a markedly different level of capability than we had a year ago, though I would call the previous two years equally useless.
How does GPT-5.6 Sol or Claude Fable 5 or Claude Opus 5 do on that Cinema 4D code?
Even 3.7. I remember when it came out and people were claiming that that was now the model that was going to replace engineers. Cue Fable years later and people still claim that this one is the one.
Yes, as a product gradually improves there will always be many people for whom version X didn't work well for them and version X+1 does. It turns out that Opus 4.5 and GPT 5.1 were larger than average improvements that cross that threshold for a significant number of people.
My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change.
If your claim is that there's no substantive difference between Sonnet 3.5 and Fable, then we live in very different worlds.
AnimalMuppet in their reply to my OP comment makes the point that these things haven't been around long enough to measure end-to-end productivity gains and come to a conclusion either way, and I agree with that. But we can still at least be measuring something.
For example, in my case AI has allowed me to write 10x more LOC than I usually would in a similar amount of time. But having to review it all, I've also deployed 1/4 the number of releases I normally would in the same period. By one measure I'm more productive, by another I'm less productive.
People could claim to be more productive by skipping the review. But in that case did AI make you more productive or did you lower standards? People could say they're using AI to do the review but is the impact of that being measured and has that caused more or fewer bugs? If more bugs, has the time to fix those been factored into overall productivity? In my experience it's common for people to eagerly count immediate productivity gains and discount long-term productivity sinks.
For this reason I think case studies are the best convincing thing, because they properly contextualize the usage and consider a longer-term window. They're also backwards looking instead of in-the-moment, so have the benefit of hindsight. But they're harder to come by and we probably won't see any meaningful case studies for a thing that people say happened in November.
But the very least people can be doing is just defining what they mean when they say "productivity" because otherwise everyone is talking past one another.
What shape did that evidence take?
The key thing is that it's easy to contrast the old way vs the new way and the evidence become obvious.
The thing with LLM tooling is that they're not reliable. I can do fine with risks, but only when there's a way to manage it so that if the disaster happens, it's practically a black swan event.
Typing more code or solving one task has never been the core problem. The core problem has always been to encode a whole system into the computer AND then provide a control interface for it. It requires both an understanding of the system you want to encode (especially how it behaves over time) and empathy to know what would be the best control interface for the users.
That understanding does not rely on the amount of code, and the best control is found through communication.
If we take the following project that you did:
https://simonwillison.net/2025/Jul/17/vibe-scraping/
An understanding of the system could be the following: A conference schedule consisting of events (time, place, speaker, description,...) stored or presented in some format. The interface would be: A web app with a mobile first UI that presents the information in an accessible manner (highlighting, filtering, exports,...).
A relatively quick (I haven't tested it), would have been to open the web inspector and extract the data using the dom API (requires knowledge of the dom api and a desktop browser), put the data into some json or a tsv file, then write a php script or a python script and then serve that. The interface could have been built with the standard elements of some css framework (bulma?).
Not saying the above is better. But the thing is that is doable from even a raspberry pi. And more it's repeatable and extensible. And the individual piece of knowledge are reusable in different situation.
Pro: "The evidence hasn't shown up in statistics yet; it's too new!"
Con: "And won't this destroy maintainability?"
Pro: "Show me the maintainability disasters caused by AI."
Con: "I can't yet; it's too new!"
Both sides are playing the "it's too new" card when asked for actual evidence to prove their claims. In fairness, it actually is too new for there to be much statistically-valid data, especially if the inflection point was November 2025. So both sides are trumpeting their position, neither with actual trustworthy data.
Everybody has their anecdote. Nobody has data yet.
I believe Aider avoided adding that for safety concerns. Claude Code demonstrated that throwing safety to the wing somehow kind of worked out.
Things are allowed to get better more than once!
The idea that "yeah, you said technology had improved in the past, and now you're saying it has improved again" is a gotcha just seems incoherent to me.
But this didn't seem to concern them. The study said X, therefore it applies to today.
I'm just not sure it's worth it to argue with others about it at this point. Not that I'm 100% all behind AI coding, but I'm just shocked people are still this resistant.
This is the correct response. Let closed-minded people do their thing. Makes them less competitive against you. There's nothing for you to gain by trying to help them understand what they are missing.
Or, put differently: if a study comes out next year saying they don't see major impacts from AI in 2026, will you admit your viewpoint as being wrong, or is your response going to be "no, there was a massive inflection point in September 2026 that completely invalidates the paper"?
> The debate over whether AI is taking people’s jobs may or may not last forever. If AI takes a lot of people’s jobs, the debate will end because one side will have clearly won. But if AI doesn’t take a lot of people’s jobs, then the debate will never be resolved, because there will be a bunch of people who will still go around saying that it’s about to take everyone’s job. Sometimes those people will find some subset of workers whose employment prospects are looking weaker than others, and claim that this is the beginning of the great AI job destruction wave. And who will be able to prove them wrong?
Or, more likely, since software touches every single industry of man, you're seeing AI slowly able to handle different types of work. People in the types of work that early models struggled with are now enabled. The goal post didn't change for the people in each group. It's not one entity.
The last people to say that their work will be impacted are those that work in areas with the most novel ideas/innovation, and/or working on things where libraries/examples don't exist.
For example, I work in test/manufacturing, mostly with robots. The whole industry is proprietary. There are basically no open source libraries/code for this stuff, so claude is still pretty terrible at it, but, with Opus 5, it is able to now do some of the work!
Everyone has an anecdote about someone being laid off because of AI, and everyone has anxiety about the next model being the big one that will really cause the industry to implode, but there's never any hard labor impact. It's always still coming, still waiting for the next model, next study, next jobs report. As Smith said, people who believe that AI is going to take jobs will never be convinced that it's not coming eventually.
You miss my point. There is no single software industry. There are people solving completely unrelated problems with software. Some of those problem spaces could be services with old models. Some only with new. It's a rise in tide, with little towns completely drowning, and the ones up the hill still safe for now.
A year ago claude could NOT write code for the robot controller we use that was usable. Now it can. The water is now ankle deep for where I am.
I didn't then create a side hustle for repeatedly re-observing that thing.
This just reads like another variation of “it’s the user not the tool,” which is just endless runway for always blaming people and never acknowledging the limitations of LLM’s.
I’d be curious to hear how the recipients of your work enabled by the “productivity multiplier” feel about the quality.
I would say that as of July 2026, with the right scaffolding, you can get reasonably good output out of a LLM, or better a combination of LLMs. For example, it pays off to prepare an implementation plan with one LLM and then let another LLM check it for flaws, then again. After several iterations like this, you will have a plan better than whatever you could come up with yourself.
It often is the user and not the tool. LLMs are complicated, have nontrivial failure modes, and the user needs to steer them carefully. They might be the most complicated tools on the planet right now.
Anecdotally, the recipients of my work have become visibly more happy in the last months. LLMs are great at diagnosing subtle problems which tend to appear at Friday night only, and this is the sort of problem that bugs actual people the most.
Totally agree, I don’t think I said or implied otherwise.
And yes can it can be the user and often even is, but when it comes to any LLM conversation I’ve been a part of it seems people think the only answer is “you’re using it wrong.” Evangelists swear it’s a 100x multiplier and anything counter to that means you’re either a Luddite who is blinded by politics or are too dumb to use the tool.
We actually have trendlines on user reported bugs, and I am happy to report they show a significant downtrend.
Of course don’t let me assume, maybe you have a higher quality disproof for the Jacobian conjecture you could share with the class.
I have had good results in that area with GPT 5.5 in the past.
On the subscription plan I don't use anything but xhigh effort and Fable, 5.6 Sol, or now also Opus 5.
Is that the class of model that struggles with Docker for you?
So whereas before my debugging time would have been spent in Rust docs looking up traits and such, these days I feel more like an AI therapist trying to figure out why it's not feeling up to task on any particular day. It can be anything from regular service outages to geopolitics that on any given day my workflow is fucked up.
In my before-AI workflow, I was never restricted from compiling Rust code because of concerns that Cargo is a national security threat. So you really have to broadly scope the notion of "reliability" with these AI tools; it's much larger than whether it can give a good output but whether it can do so consistently enough to depend on.
I've been saying this for years.
What they care about isn't software delivery, is physical goods or services that aren't related to software, for them software is a cost center.
And even then, within a single group, you'll have multiple thresholds, of "this really helps make my coding more productive" to "I no longer type code, just review" to eventually "I'm no longer employed".
The last 10% is always the hardest part, though
If model scaling holds out, we're "early-ish" in terms of the reliability and performance of these systems, just based on utilization of the compute from the planned capex. If we hit hard diminishing returns and we don't find architectural/data workarounds, that would put a wrinkle in things, but I suspect that the AI we have now is capable of helping us find those workarounds and keep things moving.