I wonder if managers will be as excited about AI when the prices go up.
I wonder if managers will be as excited about AI when the prices go up.
I suspect the API prices are already served with profitable unit economics. The SOTA API prices are much higher than the costs for other providers to run very large open weight models.
The monthly subscription plans were being offered at a discount to generate interest in these models.
We're not entering a period of billing AI at cost. We're entering a period of exploring how how the prices can go before losing too many customers.
Products and services aren't sold at cost. They're sold at the price the market will bear. It takes some experimentation to find that equilibrium point where you make more profit per customer but don't lose too many customers.
There is absolutely no evidence to support this.
I want to see their S-1s, then we can fight.
In my opinion, it’s a profitable kind of service. They probably don’t pay the public prices for the cloud GPUs though.
Or, as I would say if I were Bugs Bunny, “Duck Season”
Obviously this is an extremely rough calculation. I can even be off by a factor of 10 and it's still a pretty good return.
Training is akin to the cost of building the software/product. Inference is selling the product.
Analysts like Semi-Analysis have done a lot of modeling and estimates on the topic.
But two can play this game: There is absolutely no evidence to support that API prices do not have profitable unit economics.
Typically the burden of proof is on the one making the claim.
They have some of the best publicly available analysis on these topics. The full details and numbers are hidden behind the institutional accounts which are priced for investors (not something you sign up for personally) but they're generous with what they send out in their newsletter.
If you're not familiar with resources like this I could understand how you'd assume that the providers are hemorrhaging money on inference costs, because that is that story that gets parroted around spaces like Hacker News.
You could ignore all of that, though, and go check OpenRouter to see how much providers are selling high parameter count models. They're not entirely at the level of the SOTA models, but the biggest open weight models are not that far behind in complexity either. They're being sold an order of magnitude cheaper than what you pay for the APIs from the major players. We don't know exactly how big the major models are, but it's unlikely that they're more than 10X more compute intensive from the leaks we do have.
I see it as no different from the previous generation of consumer startups burning money - as Derek Thompson wrote,
> ...if you woke up on a Casper mattress, worked out with a Peloton, Ubered to a WeWork, ordered on DoorDash for lunch, took a Lyft home, and ordered dinner through Postmates only to realize your partner had already started on a Blue Apron meal, your household had, in one day, interacted with eight unprofitable companies that collectively lost about $15 billion in one year.
The conversation around AI being cheap now started when ChatGPT launched in 2023
> I wonder if managers will be as excited about AI when the prices go up.
Companies are willing to pay the api pricing. Engineering time is very expensive and AI coding agents actually work now since December and are actually showing measurable productivity gains, finally. It’s a good deal to make (obviously, with caveats: you need to make sure your tokens are going on productive tasks that will actually grow revenue) and anyone who penny-pinches is making a strategic mistake.
I don't think it is. At some point they have to make money and they can't do that if the token cost doesn't include ALL the costs. Someone has to pay for that at some point. And someone has to pay for the subsidized subscribers. So no. API token prices don't reflect the real price. They are still subsidized. Just in a different way.
If your company has Figma, Github, and Cursor and they're using the same models you are, your monthly costs with them increase as well. You're exposed N times to the foundation model price increases, where N is the number of times software you directly or indirectly use talks to a frontier model.
I always wondered about this statement, like we are generally salaried and there is so many variables that affect how I spend my "time". None of us are machines that can do X work per day and our managers get to slice it as they see fit. Pull a dev off a project they love and throw them onto something they hate and suddenly X is diminished greatly.
I would almost predict that reshaping our workflow to be: "prompt, wait, approve changes." results in losses because it is such a mentally tiring workflow and drills into our brains the desire for the LLM to "just fix it". It is the next level of just moving tickets to completed all day.
Citation needed. Anthropic does not have public books
If their CEO was just flapping his mouth without any other comparable baseline, it'd probably be different. But as the GP points out, open-weight model providers are charging comparable rates and very likely have positive profit margins. That would imply that with API pricing tokens are sold at above cost.
That cost may well be "inference only", so excludes everything apart from hardware and power. Whether that's enough to cover the enormous training costs and other overheads is a different question.
he has access to the real numbers and a legal risk from lying publicly. it does him no good to lie about this.
It is? If another company comes out with a better model tomorrow and offers it at the same price Anthropic charges for Opus, they’re going to lose customers fast. They have to keep training to keep selling inference.
Most businesses factor in the cost of making their product into the product’s P&L.
Lastly, theyll realize like every good capitalist, theres more profit in exclusiveness and cutiing out customers.
That is no longer a helpful tool... it costs like ~15% of an actual dev.
Even if it is helping, is it actually... making things better or building anything truly important? The issue seems way too nuanced to spend $2k/mo. Not to mention the entire tech industry floats on hype and imaginary goal posts so now what? Devs can hurdle towards those faster and more mindlessly?
The full cost of each employee is more than their salary. The common estimate is 1.4X their salary due to all of the employer-paid taxes, benefits, and other things.
So even $2K/month of token costs would only be around 10% of the cost of a mid-range developer cost.
It doesn't have to increase productivity much to justify the cost.
Another challenge for US tech companies is that - if you'll forgive the bluntness - their "brand" is now toxic in most of the world. Almost everyone is trying to distance themselves from US tech as fast as they can. Governments and big businesses are starting to invest seriously in alternative solutions and local resources. It will happen over time but I don't see much the US tech companies or the US government can do to stop the train now the wheels are turning.
So there's a serious risk for US tech companies now of a double whammy where their already relatively high R&D costs increase even further and yet they're also facing much stronger competition in international markets or maybe even excluded from some of those markets entirely.
If we also reach the seemingly inevitable point that "capable enough" LLMs can run locally - or at least as a private resource provided internally by large organisations - there is very little moat left to protect not only US Big Tech whose stocks have been heavily driven by expected returns from AI but the whole US tech industry that is banking on productivity gains from that AI tech. Then they also won't be able to capture most of the entire global supply of components like GPUs/RAM/SSDs because it won't be cost effective any more - and that is one of the few practical moats they have built (however accidentally) that would be a significant barrier to direct competitors setting up shop in places like Europe and Asia.
It's going to be interesting to see how US tech companies respond to these effects over the next 5-10 years. The giants are all aboard the AI train and can't back down now so there will probably be some casualties there if - as again seems inevitable - the bubble bursts at some point. But then there's a very long tail of still very successful US tech companies that might be paying US salaries and using AI-based tools but aren't themselves focussed on developing or providing those AI-based tools and they're the ones who are going to need to find new ways to compete effectively within that kind of time frame.
What they don't like is paying money for the work, that's all that matters to them.
Thus, your compute is significantly more expensive than AI. Thankfully your taste is also part of your package deal, and is where you deliver real the value over an LLM.
IMO the programming world is far too myopic about / insistent on using laptops, especially macbooks. Just because a crappy deal exists doesn't mean everyone is forced to take it. Local AI is a high performance computing problem and laptops are fundamentally a crappy form factor for it; buy an efficient desktop computer and be surprised at what's possible even with today's crazy prices.
I have used the following on a 32G MacMini to help write useful code:
ollama launch claude --model qwen3.6:27b-coding-nvfp4
The problem is that running local models (except for engineering tasks like data munging) is slow. With the above setup I set up a task (asking for no user verification) and go for a walk to wait for results that my Gemini Ultra plan would produce in 10 seconds.