At this point I'm trying to believe there's a middle ground where the level of individual capability this unlocks, leads to major discoveries.
At this point I'm trying to believe there's a middle ground where the level of individual capability this unlocks, leads to major discoveries.
Take any stock index, remove AI stocks, what do you see? That's right! Nothing...
So where is all the productivity going? Where is the value? Where are the massive unemployment stats or the millions of new startups making big $$$?
That being said, AI seems kind of miraculous sometimes.
Similar to cars. So enticing that we make everything else in the world worse in order to maximize the profit, make it indispensable, subsidize it, and make the dependency on it irreversible.
And it's not even something to blame individual people for.
Driving away from all the other cars to spend a weekend feels like freedom.
Using AI to answer a question feels like a "bicycle for the mind".
But in fact it's more like a car. It requires massive resources and creates perverse incentives, and the result is ineffective and corrupt.
Both cars and AI are amazing technology and extremely useful, but using them is not an individual responsibility. It requires societal subsidy.
We got addicted to the convenience and overuse, and have started a mass extinction event because of it.
The perverse incentives will come for us all.
It feels depressing, but I think the same. When thinking about the larger world, it becomes increasingly hard to ignore. And of course it is not new.
There were "doomers" already in the midst of the 20th century, but it doesn't mean that they were wrong.
AI, Crypto, car-centric urban planning, cardboard suburbs in a barren field with bright green grass.
You tell someone that their selfish choice to drive a pickup truck only to haul something maybe one time a year, they shit their pants. They can't stand it, they start personal attacks and ask "what about you??". They seem to know it's wasteful, damaging, etc but they deserve it for some reason.
80% of generative AI queries wouldn't even exist as google searches.
I awe at the capabilites of generative AI.
I also enjoy sitting in or driving a car.
I did not want to make a moral argument, unless you consider each and every form of utilitarianism as moralism.
One claim of the parent comment was that AI is ineffective. For the purpose of finding answers to questions, it is more resource-efficient than the alternatives, and, to your point, capable of answering questions that were impossible to answer via other means before. In what way is that ineffective?
If we want to get really pedantic, every generated token is the answer to the query: what's the next most probable token in this sequence of tokens?
When I post this http request containing this reply, you could say my machine is querying the server to ask "what did you do with the message I just gave you", but then query stops having any useful semantic value to distinguish it from "request"
Regardless, this is tangential. I don't disagree that a lot of LLM use is not in pursuit of knowledge, but enough of it is for me to think that preferring LLMs not to exist is a hard position to defend, at least without making the case for existential doom.
Neither did I want to say that a car is always more wasteful than some alternative.
But defaulting to the behemoth is inefficient, unless everyone is driven to do it: then it's in some way reasonable.
By adding "corrupt" and "dependent", as well as the economic terms, I wanted to offer a broader critique and create an analogy, not just talk about energy usage on its own.
What I had in mind was: it's easier to go many places that are a mile or less from me, by car. Because everything is obstructed by cars. And I'm atrophied by lack of movement. Best would be to drive somewhere to move/walk.
People already do that in masses.
And doing shopping by car, because everything else seems unbearable, also takes away your time, apart from wasting energy compared to more, smaller shops that would be reachable by foot, bycicle etc.
I guess you know the argument.
Today, people's thinking atrophies because their LLM is probably right in their summarization of some Wikipedia article, plus 2-3 other random sources.
Or so.
Using the Wikipedia search function is not expensive.
But, I mostly had a bigger picture in mind than what is the cost of inference.
I am concerned about the environmental impacts that AI poses, but they don't seem to me to be so catastrophic. Solar and battery tech has made enormous leaps in the past couple decades, and we will need to pivot to clean energy future irrespective of AI.
*This said, I have become gradually more alarmed over the past decade at the lack of epistemological rigor in the general public, as made apparent through the rise of social media. I don't know that AI becoming a truth-seeking crutch for people wouldn't be more good than bad.
Oh my god, no. I also want the benefits of automobiles! They are strictly more capable than, say, trains. That's where I would derail the discussion completely when going into details, but no, I am not against cars as a technology.
Apart from all the ethical and social arguments (logistics, ambulances, the elderly, etc etc). But that's not where I wanted to go.
I was making a leap here simply because of the whole complex around prisoner's, dilemma, the commons, state economy, and so forth.
Since at least ~100yrs ago, I guess cars and streets as the primary mode of transportation have also "won the vote" / are what the majority wants, so it's also an interesting analogy for diminishing returns maybe.
Building out more car infrastructure is certainly not controversial where there is absolutely none but there are commercial or residential buildings.
Anyway, lots of associations are worth considering here IMO. The ultimate limiting capacity here, when disregarding all environmental or health concerns, is simply space and the positive externalities (cities etc) around existing infrastructure.
> Take any stock index, remove AI stocks, what do you see? That's right! Nothing...
> we make everything else in the world worse in order to maximize the profit
> destroying the planet for data centers
Hand wringing about AI datacenter's environmental impact is well and good. We should keep the data centers accountable for their consumption and waste.
I just wish the same people had been upset the last 20 years with poor water resource management in a lot of areas (the west US especially) with urban, ranching and farming development.
> That's true, and I am not anti-AI.
Me neither!
But with AI what is the exact price? My understanding is that R&D is extremely expensive, but running non-SOTA models is not that bad. We are getting pretty close to models which can be useful locally in many applications.
Or do you mean that at scale running them locally is not possible and hence the infrastructure price is in data centers, which will be expensive to maintain and scale for demand?
First, because I initially failed to answer your more closed questions (this paragraph is edited in):
> We are getting pretty close to models which can be useful locally in many applications. Or do you mean that at scale running them locally is not possible and hence the infrastructure price is in data centers, which will be expensive to maintain and scale for demand?
I don't think there's a way around making the best of AI capabilities with minimum price and maximum control, and I'd agree this is met by on-prem data centers, just not in a rationally targeted way.
Back to my original comment:
Because it (my conclusion) was not so clear, and maybe I just wanted to highlight some observations without delivering a real argument for or against things [, I thank you for your open question].
The utility/leverage aspect for AI seems more esoteric than the one for cars because, apart from Chatbots, it's more hidden.
And also, similar to cars (or many other phenomena of industrialization), yes, my first vague point was the subsidization of infrastructure. But also, the power gap: that's something not only associated with AI or cars, but with a lot of technologies we all hold dear: sewage, powerline, logistics, etc etc.
What reminds me of cars in the current AI frenzy is the fixation on cementing infrastructure. And also, I think, a lot more people agree on, for example, some kind of universal right to, for example, clean water.
But all of industrialization confronts people with questions of efficiency, inequality, and collective support.
Most people would, for example, support a right to get a minimum amount of clean water when you are living and working in a tradionally inhabited space (if you're on the social-darwinist side) or at least not harming society (if you're more of a social democrat).
And, similar to the buildup of car infrastructure, and the procurement of resources, space etc for maximum building, giant data centers can obstruct people in buying drinking water. Or walking outside (AI obstructs traditional methods of online collaboration).
Where did all the stock gains go before AI?
FAANG / MAG-7.
Was everything from 2012-2020 fake, too?
The question is, is AI leading to massive productivity gains in companies that implement it? AI productivity gains take time to diffuse, but so far companies in the S&P 500 are seeing very high growth. YOY earnings growth rate for the S&P 500 is 21.7% https://advantage.factset.com/hubfs/Website/Resources%20Sect...
Now remove the companies selling the AI shovels: https://pbs.twimg.com/media/HIAjbZxacAARHwD.png
> Not sure what your point is.
My point is that they're selling us Skynet and the end of employment as we now it, things that we shouldn't even have to measure to perceive the results of, yet no one is able to measure any of it
Pointing a finger at nvidia, google, and the other few companies stuck in circular investment schemes that shouldn't even be legal and saying "OOGA BOOGA line go UP, UP GOOD!" doesn't count in my book
https://insights.som.yale.edu/insights/this-is-how-the-ai-bu...
> AI-related stocks have accounted for 75% of S&P 500 returns, 80% of earnings growth and 90% of capital spending growth since ChatGPT launched in November 2022.
If all these false practices can pull revenue out of nothing, why doesn’t every company do it? How come AI companies seem to be able to pull off financial magic that no other company can match?
All your analyses still ignore the revenue point.
Then why can't anyone point at actual numbers? The best we get is "look: line go up" while pointing at either the companies selling the AI shovels or the companies selling $1 of tokens for 50ct.
When cars replaced horses we didn't have to twist the numbers to understand the benefits. When emails replaced mails we didn't have to do 6 hours of mental gymnastics to see the increased productivity. Heck my grandma could tell the benefits of computers when the town hall she worked for finally discontinued typewriters
They're spending hundreds of billions if not trillions, and have nothing to show for it besides like 5 stocks pumping like shit coins. On top of the the drawbacks are massive and very visible...
> Take any stock index, remove AI stocks, what do you see? That's right! Nothing...
Parent comment:
> Now remove the companies selling the AI shovels: https://pbs.twimg....
From your linked image, "excluding AI stocks" is "+16%" (the figure with AI stocks is far higher).
Your sole source says +16% excluding AI - in what kind of market is +16% “nothing”?
It's nothing because it happens all the time, it's not statically relevant, like not at all: https://www.macrotrends.net/2526/sp-500-historical-annual-re...
This forum is full of techies with very strong opinions about their toys but 0 economical, political or historical education, and it shows
And even more so since inflation was 2-3%, not considered high, during most of that period.
I mean, do you know what the value of those stocks would be if AI didn't exist. Maybe they would be much more negative. Maybe we would be in a recession. Without a control this type of analysis is meaningless.
And that is even assuming that AI productivity gains are happening now instead of 5-10 years from now.
Infrastructure doesn't produce value overnight. How long did it take the Interstate System to provide measurable value? I asked Gemini. Supposedly increased national productivity by 25% over 39 years[1]. But if you drove on a newly finished interstate in 1959, you saw the same cars just moving a lot faster.
That's what we're seeing right now. People can produce an incredible amount of stuff really quickly with AI. Is it directly connected to measurable productivity across the entire economy? No, because, realizing a mass productivity increase from infrastructure takes time.
[1] - https://www.richmondfed.org/publications/research/econ_focus...
I do value having some naysayers in the mix generally, because we do need balanced critique in what is otherwise a very frothy hype cycle. I just don't think he's making sound arguments, and that's even assuming you even agree with his premises in the first place.
My biggest gripe with his napkin math is that he treats inference gross margins as something novel that you can't compare to normal SaaS margins. He's right in part: the constant carousel of R&D costs from model training, related infrastructure buildout, and other adjacent costs required to stay competitive do change the analysis a bit.
But he takes this way too far when he says this is structurally different from normal SaaS margins. The business model definitely doesn't look like Dropbox, but it absolutely looks a lot like AWS, especially early AWS, CDNs, telecom, etc. I can speak to the telecom bit personally, since it's been over half of my professional career as an engineer and, in this specific case, also as a founder. You can have a brutally capital-intensive infra business where profitability depends on utilization, oversubscription, peak-capacity planning, segmentation, and recovering capex over time.
The math he presents gets even more questionable as we see explicit segmentation happening for cost-saving reasons. Many forward-thinking orgs are waking up to the fact that they don't need to use the best, most expensive model for every task. They can route easier tasks to cheaper models, use caching, batch non-urgent workloads, and reserve frontier models for the subset of work that actually needs frontier intelligence. That directly undermines his claim that providers always need to chase frontier intelligence in order to maintain current demand, utilization, and pricing curves.
Could you share what tells about it? I.e. where he was wrong about it?
I'll cherry pick a couple:
“When these new models ‘reason,’ they break a user’s input and break into component parts, then run inference on each one of those parts.” [1]
This is not at all how test-time compute works. At best, this is a very loose metaphor that he may have used out of convenience. This might sound a bit pedantic to point out, but this is a very basic thing that he's getting wrong (presumably at least, again it could be that he just used a poor metaphor).
A less pedantic example would be his claims related to gpt-5/chatgpt auto-routing. He argued that having a router means OpenAI can no longer cache static prompts, because the user prompt has to come before the hidden instructions [2]. This is just not at all how this works at inference-time. There is no evidence that the standard approach of system>developer>user instruction hierarchy has changed, the public API and caching docs maintain this.
But even more broadly, it suggests he is reasoning about kv/prefix caching at the wrong level of abstraction. It's true that conventional prefix caching does require a stable prefix, so yes, if you literally put variable user content before the static prompt, you would destroy the cacheability of that static prompt.
But that is exactly why inference systems are designed to preserve reusable prefixes where possible (via checkpointing or similar), and why serving systems care so much about prefix caching. This is also a big part of how disaggregated prefill/decode infra works where cache-aware routing is critical. His argument treats a bad prompt layout as if it were a necessary consequence of routing, rather than an avoidable implementation choice.
A router can read the user request, decide which model path to use, and then construct a normal downstream model call with stable static instructions first and user content later. Treating that as impossible implies a fundamental architectural misunderstanding.
[1] https://www.wheresyoured.at/how-to-argue-with-an-ai-booster/
But that is not the full argument he is making. If the claim is that the labs will not be able to pay their creditors because inference is structurally incapable of becoming profitable, then he absolutely needs to be right about the technical economics of inference.
One part of that is the balance-sheet argument (which already shows insanely good margins). But it also depends on how inference-time compute actually works: routing, batching, kv cache reuse, model segmentation, different latency tiers, etc. Much of those details he's just been straight up wrong about in his writing, so as a result I have to call into question the rest of his reasoning as well (in part to avoid Gell-Mann amnesia).
Also, if there is significant gains from caching, then like.. what are even doing here? Inputting something and then reading cached pieces of text based on their similarity to the input? Kinda like a search engine?
The newest biggest model can still matter even if you do not run every prompt through it. You'll always have some task where even small amounts of loss are unacceptable and thus you need to make sure frontier intelligence is used for it.
On the router point, yes, routing has some overhead. But the router does not need to run the biggest model to decide which model to use. We've been using tiny classifiers for recommendation engines for ages now, usually on CPU. If routing saves you from sending a large fraction of traffic to the expensive reasoning model, the routing overhead can easily be worth it.
> Also, if there is significant gains from caching, then like.. what are even doing here? Inputting something and then reading cached pieces of text based on their similarity to the input? Kinda like a search engine?
The caching I'm talking about is explicitly the attention/kv cache, so its not input similarity retrieval (that would be more like what you'd use in a RAG/IR system). Prompt caching is generally about reusing already-computed attention scores for repeated prompt prefixes. The idea being you don't recompute the same static system prompt, tool definitions, schemas, long shared context, or repeated boilerplate every time. In more sophisticated systems, you usually store multiple checkpoints so that a small prompt change doesn't result in all-or-nothing hit/miss scenario.
But does it also not mean that they will make less money given that there is already brutal competition for that lower tier from openrouter, Deepseek, Amazon, etc.?
You can't on the one hand say "customers are beginning to understand they can spend less" and on the other hand suggest that this is good for forecasts of revenue.
Sure you can. Just because there is a non-zero amount of margin pressure from the lower tier inference providers does not imply that revenue forecasts ought to be poor. Jevon's Paradox gets oversold in this current cycle, but I do think it's a relevant lens to view this through given how much demand has outpaced capacity.
The argument is that customers learning to spend less per task can be good for the viability of the market (really the total demand) even if it is bad for naive revenue-per-token assumptions. If a workflow goes from economically stupid to economically viable because you route 80% of it to cheaper models and reserve frontier models for the hard cases, that can expand total usage and improve cost per useful outcome.
However, most of the engineers I respect have gone from being skeptics a year ago to convinced today. I don’t personally know any true holdouts any more. If there are studies that disprove productivity gains more than six months ago, I’m happy to believe that it was true of the AIs that were available at the time. But I’m going to need something much more recent before I disbelieve my lyin’ eyes where it pertains to the AIs available today.
Here is the report:
https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
And my commentary:
My view is that it's not really about how good the models are - it's about how we're using them. Understanding what you've built is an important part of value creation, and LLMs eliminate that.
I currently don't have work access to Claude Code, but most of my teammates do. Watching from the outside, the cycle seems to look like this:
1. Experience some success, which hooks you into relying on AI.
2. The AI keeps failing at some task, but you don't want to stop. Keep trying over and over again.
3. Run out of tokens and take a break.
Now, sometimes 1 doesn't happen. Sometimes 2 doesn't happen. 3 is a certainty though.
Now, if you told me that the productivity gain from 1 is enough to offset the loss from 2 and 3, I could believe you. But I also wouldn't be surprised if it didn't.
Even if you take for granted that AI is as good as the best people say in writing code. And Ive spent a lot of time generating codes, I won’t disagree - Then the question becomes - does this change your daily incentives such that you reach for code as the solution to your problems rather than something else (coordinating with your colleagues? Product management? Planning and Design?
So from a holistic perspective, I think intentionally limiting your own AI usage is the best approach for maximum long-term productivity.
But what if the problem you’re trying to solve is the altogether too often problem of like getting teams that are dependent on you to upgrade the library they use. And what if the library is a breaking change, and last year they upgraded to the library on your advice and it broke production and now they’re suss and want to accept all changes, and integrating that library change isn’t in their critical path so they’re just not going to spend time on it, even if you submit the MR them. Even if you show them their tests pass after the change.
Importantly to the above, you probably need more devs to do more of the above in parallel. You don’t hire devs to write more code, you hire more devs to carry on the mental load of a broader scope of work. Even in the before times, so much code got stuck at the integration step.
But because all that is hard, instead you go and codegen to fix an obscure bug that sure makes a few customers happy, but no one thought was a limiting factor for paying your company more money.
It’s not that I don’t think AI can help, I think it’s a prerequisite for the job and everyone should use it. It’s more that I think in the grand scheme of things, people will bias towards using it for tasks that aren’t in the critical path - refactors, tech debt, bug smashing, tool building; and I think it could really help devex and that’s good.
But I think people are bad at knowing the difference between “my job feels a bit easier” or “I’m more productive” and “this task had an impact on the bottom line” and when you extrapolate that out to a whole engineering org, that’s where the productivity statistics get lost.
I’ll addd one data point to this is like this thread itself. So many people on AI skepticism threads point to their own subjective experience as evidence we’re not in a bubble, and sort of ignore the entire concept of economics. I’m not saying we’re in 100% in a bubble, but subjective experience isn’t great evidence of it.
And this is just sort of one of the factors, what about the increased cost and mental load of supporting more software? What about junior engineers who feel pressured to ship work but don’t actually learn the software engineering? What about lost context from not intimately understanding your software?
Although if this theory is true — that AI helps with coding but coding is not the friction point in organizations with multiple humans, even that should allow faster iteration by allowing one human to do more coding therefore reducing the size of teams required to make some programs. You should see good acceleration in solo shops too.
I’m a platform engineer. The primary failure mode for platform engineers is building tools people don’t want. AI doesn’t really make that easier. Or it can but it can also make it harder by making it easier to chase down ideas that you don’t get traction on. And I think that - net balanced across the organization is probably why productivity gains get sort of averaged out.
For sure I think solo devs who have a system are seeing gains, as long as you can I think have the discipline to have a process that includes feedback and learning and your not just feeding off of dopamine hits of one shotting features but yeah. I mean for solo devs the code was never really the limiting factor, it was product-market fit and marketing.
So solo devs who have a system may be laughing themselves all the way to the bank, but we may not see a lot of net new solo devs.
But if code is cheap now then it’s sort of inherently devalued. 2 8 person startups can probably relatively easily find a dev with AI experience to rocket ship their code generation, which means the basic skills of talking to customers, change management, and building the right thing become even more valuable.
Even solo devs I wonder - almost every post-mortem of a failed company goes “I wish we had spent more time talking to customers and less time writing code”
Again if you can get the discipline right, maybe as a solo devs you can get more work done faster and spend more time with your family. That’s incredibly valuable!
But if you go and add a big new feature, or a second product - unless your community is primed for constant growth(no man’s sky is one community where more more more seems good) you’re just growing the surface area where all the other skills are more necessary.
I think this is right. They are much better applied as editors than authors, IMO.
The key thing is stay in control of your output. i.e. understand it thoroguhly. I think you let the LLM make decisions you don't really understand, you're increasing the likelihood of introducing defects that are expensive to address.
EDIT: In fact, parent comment has a link to some numbers.
[EDIT: Most] people don't want to go through the numbers. Ok. But there's a history here. When people don't want to see the numbers, certain kinds of things tend to happen.
Code acceleration is great, but.... something precedes that. Vision and strategy re. expansion of offerings and businesses. Once a firm reaches maturity in what it offers and is only touching the edges - this code acceleration is literally useless when you factor in all of the trade-offs.
This is a good thing - it means fat and slow incumbents are sitting ducks to be out-witted by creative and imaginative founders, which is healthy for a well-functioning economy.
Now the economics of existing frontier models are not sustainable - its looking like a mix of the airline (supersonic vs subsonic) and EV industry with China in the background providing decent offerings at much lower prices.
I admit that if a small team or an individual uses an LLM, it's likely they can create value faster.
I think as soon as you don't own the responsibility for the defects you generate with an LLM, their use starts to destroy value. Regardless of product maturity.
This is what I think the data says.
I actually think this is precisely the reason LLMs can't be the basis for a technological revolution. Because it's only one way.
Like, if you have a compiler, and it has a bug. You can discover if that bug is influencing your code execution and patch it. You can go both up and down the stack.
With LLMs, there is no way to patch it's translation function. You have to rely on it to forward process.
I don't think there is any way to avoid us understanding our tech stacks.
If you are producing something that delivers a far better experience, irrespective of what's under the hood (see Claude Code et al), you will decimate an incumbent who is trying to use LLMs in the context of incrementally improving a mature product.
LLMs are suited for the development of revolutionary innovation, not incremental.
I think I just disagree about the power of the LLM to deliver revolutionary innovation. That's something you do. Not the machine.
And, pretty soon on your journey to scale, the LLM becomes a hinderance rather than a help.
The thing people I think have a hard time seeing is that "I go faster" does not mean "more features get finished".
It's a scale issue, and one scale is better than the other. People only pay for finished features, they do not pay for how much code you emit.
In my field - operations - productivity is usually described as some rate of production for a specific asset. 100 widgets / machine / hour - for example.
"My productivity is 3 PRs / day with the LLM as opposed to 1 PR per every three days". That's how I think people are thinking about it.
My point is that's not the same thing as value. I.e. what people will pay for.
“This random part of the code is slow, I used an LLM to generate a PR that speeds it up.”
Okay, you optimized the part that’s not a bottleneck, sped up nothing and cost the company $100 in tokens. Good job?
"If an LLM builds a feature, and no one uses it, did it make value?"
You're right my analysis is at variance to what Faros.ai says. I think they interpret their data trying to rescue utility for the dominant patterns of LLM use.
But I think to anyone who is experienced with process improvement or queuing theory, their interpretation is clearly weak. Rework is a huge problem in queue systems, and they mostly just elide the throughput impact of an 860% increase in code churn coupled to a massive spike in bugs.
Obviously draw your own conclusions. But I don't think because I disagree with the interpretation of the people who originated the data makes me wrong.
The fact he’s never reflected on the glaring failures in his analysis tells what we need to know about his intellectual integrity. There’s truth in some of his words about financial risk, but if you can’t acknowledge that there’s upside too, you can’t evaluate risk properly either.
I find it difficult to take him seriously.
Do you think it's not slowing? Do I miss anything really important?
My understanding is that we have now is incremental improvement on thinking models which appeared more than a year ago. Of course, a breakthrough might happen, but I don't see one yet.
I think it's dangerous to rely on claims made by people who financially profit from you believing them without checking.
[0]: https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-v...
And presumably the GP thought that saying the maintainer had access to Mythos made it a more compelling argument. Otherwise why even mention it?
It found hundreds of vulnerabilities in Firefox, according to Mozilla: how does Mozilla benefit? It found a 27 year old vulnerability in OpenBSD. How do they benefit from that? Is that made up? Are the maintainers of those codebases lying for the benefit of Anthropic’s IPO? Is copy fail a fabrication by big AI? The 12 OpenSSL vulnerabilities found in January?
https://venturebeat.com/security/mythos-detection-ceiling-se... https://www.wired.com/story/mozilla-used-anthropics-mythos-t... https://cyberscoop.com/copy-fail-linux-vulnerability-artific... https://www.schneier.com/blog/archives/2026/02/ai-found-twel...
Im not sure whose claims you think I’m relying on. I trust Firefox that they’re not overstating the number of CVES they’ve found. Same for OpenSSL. The OpenBSD folks definitely don’t seem like the types. I’ve not known Linux to fabricate CVEs either. I think my sources are fine.
Have a muck about with what Qwen 3.6 or Gemma 4 can do and you'll see. I mean this as an illustration but Qwen just isn't as far behind as I expected, and compared to the data centre hardware it will run on a potato.
The frontier models are losing their undeniable edge over that which is unmetered.
And even putting aside my optimism for the smaller open weights models, there's a huge amount of scope for the larger, hosted open weights models that are only just behind the cutting edge and which cost, what, 1/25th of the price on opencode go, openrouter etc.
Commodification is coming, and with it slimmer profit margins; it's hard to see them making anywhere near the kind of money they need to in a commodified market.
Old WSB saying: The market can remain irrational for (far) longer than one can remain solvent.
And unfortunately, a lot of the market on the "buyer" side has been acting irrationally. When you see CEOs telling their employees that they don't care about token cost, only about "how much AI do you use" because that is what the stock market wants to hear - that's when you know we're all getting cooked, the question is how long it takes until the bubble bursts.
It's not that the utility of it put in question. What is however a giant question mark is how the heck any of the big AI companies are ever gonna get that ROI? Given how many of us are becoming more and more fine with local models that run just fine especially on a good enough computer which most developers have anyway...
Why should someone pick Opus 4.8 when Qwen3.7 Plus produces similar results for about 1/20th the cost.
That sort of pricing disparity is across the board. But further it's becoming more and more apparent that they are doing more with less parameters. That's what's giving the local models their super powers.
I'd say that yes, ignorance plays a role here because a decent number of engineers are looking strictly at the benchmarks and choosing Opus just for that reason.
But I'd also say that a major factor for Opus use is because Opus is being purchased for the engineers by their employers. They don't get to pick which models they are using.
How can something so undeniable have zero scientific evidence? Are there any large peer reviewed or meta studies confirming your claim?
I think the surest sign of productivity gains is the sheer volume of adoption. If you look beyond headlines, adoption is just incredible. Of course adoption does not necessarily point to productivity gains, but if this was some sort of FOMO or smoke and mirrors you would not see this much retention and this feverish a pace of adoption. You would not see a large segment of the profession using coding agents exclusively. All of these companies track productivity, again with imperfect proxies, yet everything points to a pretty consistent picture. Same with benchmarks, again a lot of crappy benchmarks but a lot of high quality ones too and a very diverse collection of tasks and capabilities they probe.
Adoption meaning productivity supposes there are no other dominant factors for the AI push nor AI retention. It is possible for practices to be picked up or continued in spite of causing productivity DROPS. What studies have suggested are factors that make for productive work environments and what is actually enforced in the workplace are different things.
Adoption implying at least some significant productivity gains doesn’t contradict there being other factors. You’re seeing entire companies reshaped. The argument is this is all for show or CEOs are in some sort of idiot class?
“It is possible for practices to be picked up or continued in spite of causing productivity drops” well of course. I just find that incredibly far away from Occam’s razor.
My point is: we have lots of evidence that’s highly consistent with real productivity gains, and I don’t see many pieces of evidence to the contrary.
I agree that there are incentives to waste tokens but this is strategic: they spend now for people to build and explore, then they keep the things that work and drop the things that don’t.
LoC: people argue it’s not what’s important
PRs/day: same as LoC
Getting projects done faster: oh but what about the quality.
Solve the technical problems and actually be more productive, the social systems build around the old way of doing things will hole you back.
Finish a PR in 10 minutes doesn’t matter if you’re waiting days for a human review.
Why cant it naturally grow and prove it's worth?
And where are those? They seem particularly hard to actually observe and only appear in anecdotes.
> I'm trying to believe
For every exponential increase in compute capacity you see linear gains in output accuracy. This is a death spiral. Anyways, you see "massive productivity gains" so why is "belief" a function of your viewpoint?
The way you make a viable service that eats 300bn annually is to have enough demand to service that. Anthropic underbought compute. That tells you something.
How far behind are models that can be run locally, and do you expect that this will be widespread?
I think over the years local models have fairly consistently been ~7 months behind frontier performance. Local models are hugely important but I don’t see the calculus changing. I can imagine it’s certainly the case for many tasks that there will be diminishing gains for performance improvements or reliability pass some threshold, in which case you don’t need frontier performance and you can certainly use local models or at least cheaper tiers of proprietary models if local is too much of a hassle. Plus of course use cases where local is necessary or the pros of having local models or on device models outweighs that of frontier.
Maybe things will change though, I would assume through basically government subsidies from China etc, to undercut existing frontier labs, but you can always spend more (better data more compute etc) for better performance and that I can imagine will always have a selling point.
The jury is still out on that.
Uber, for example, is so unclear there is any ROI, they are cutting their exposure pretty radically.
He points out that one single Anthropic customer — a payments provider — accidentally had to pay Anthropic $500M for one month of token spend.
That is half what Apple is reportedly paying Google for the supply side of their entire consumer AI strategy.
Owning a chunk is pretty directly how more than one country injected confidence into their at-risk banks; it’s certainly how it was done in the UK.
The question is: what does "underdeliver" mean here? the pro-AI arguments I am seeing in this thread are equating mass adoption to agentic coding. Er, I dont know of any trillion dollar cap companies that sell dev tools. The point is Zitron doesn't have to be 100% right for his central prediction to come true.
* robotics (need to close data gap and release first viable product to get a data flywheel)
* conversational ai (no one is ready for this and we’re getting closer and closer to natural speech. The quality still isn’t good enough but it’ll be soon).
* other agentic use cases, openclaw adoption was crazy and that had a ton of barriers to entry
* ai products, like the one OpenAI is working on with Johnny Ive
Anyone thinking it’s unreasonable to hit whatever revenue requirements is just not that aware of what’s happening. Not to mention were capacity constrained already!! This is barely speculation at this point.
- RL is extraordinarily sample-inefficient.
- distribution shift/catastrophic forgetting aren't solved. only off-policy learning with giant decorrelated batches works.
- the breakout success of transformers as an architecture doesn't neatly translate to robot motion policy models.
the field is missing fundamental breakthroughs.
I also find it very interesting that conversational AI has taken this long. where are the models with good turn-taking? passive listening? the ability not to respond in paragraphs? has Anthropic simply not gotten around to it?
For conversational AI these labs do have lots of things to do lol but you’re right; it likely also requires some architectural improvements but you see the infancy: look at the llama4 speech duplex model. Very unimpressive yet all of the components are there. Just a matter of pushing on them, licensing and commissioning better data, etc. takes time and compute is stretched thin.
This, combined with his extreme ignorance, makes him unreadable. The only reason people read his stuff is because it validates and confirms their own anti-AI beliefs. It's why every time he publishes an article, it reaches the front page in an hour or less.
Extreme ignorance?
How are they undeniable? They're very deniable. One example is the (seemingly) increasing maintenance costs for AI-generated code[1]. Another is the cost incurred by everybody reading AI slop instead of actual communication.
I don't have hard data as to whether these cancel out the benefits, but it's not as rosy as some seem to think.
[1] After years of people understanding that LOC is not only a poor productivity metric but also a negative indicator of code quality (shorter code for the same thing is better), we now have people touting how many LOC their LLM agent is generating. It's like everyone forgot what LOC actually represents and what it means for long term maintenance costs.
No, he's not, he's making tons of money every month from his Substack subscriptions. In fact, the AI bubble popping would be the worse thing ever for him, he would be out of a job.
Just like the who have predicated the US dollar will collapse any-moment-now and which pushed gold for decades.
Funny how people always say "oh, you are an AI lab, of course you are going to hype AI", but never "oh, you make sooo much money from predicting the collapse of the AI bubble..."
Just because you keep repeating something doesn't make it an undeniable truth.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.