I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.
Imagine the power company being able to decide how much you consume and at which price point.
Airlines, banks, health insurance…
> Step 2: Regulatory Capture
is being worked towards.
Are they though? Or is it just what some companies would want them to be?
I don't know if that's the actual origin of the term nerf, but it was the first time I'd heard it.
When you're running something at a loss, you can mistreat your customers and they'll still stick around (I'm an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.
My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they'll either drop the subscription model or it'll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.
It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible, even if they're delivering real business value (coding, research, etc.). In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don't know, we can only speculate.
> Many say that, but I sincerely doubt they'd actually follow through.
I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I'll be using it less.
Look, I am not saying that there isn't a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don't have a business model here.
You can assume Openrouter open-model providers serve at or above margin, because there's no branding so there's no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they're very comfortably profitable on API.
That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won't be fast and won't be frontier intelligence, but definitely useful.
It's meant for their test env I think, so is not documented to my knowledge
https://code.claude.com/docs/en/monitoring-usage#usage-monit...
thanks for correcting me on that regard
Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?
I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.