12 karma · joined January 2, 2025
They use these releases to get users. When they got them they can play around with "degrading" model just enough not to loose users but save on costs. It sadly kinda makes sense...
But that is huge and expensive project. Only "approximation" I could pull of reasonably to get this started was to use benchmark scores as "surrogate" for that.
But working on a way to get this going. If you have additional thoughts on how to approach this I it would be super valuable.
Part of it tracks how many tokens you actually get from various subscriptions, over time.
Past week, multiple people asked me about it — they'd been hitting Claude and Codex limits faster than expected.
Ran the tests yesterday. Reran today. Here's what came back: ▸ ChatGPT Plus / GPT-5.5: 95M → 37M tokens/week (−61%) ▸ Claude Max 20× / Sonnet 4.6: 388M → 214M (−45%) ▸ Claude Max 20× / Opus 4.7: 248M → 162M (−35%) ▸ Claude Pro / Sonnet 4.6: 19.6M → 11.4M (−42%) ▸ Claude Pro / Opus 4.7: 15.6M → 10.2M (−35%)
5 of 5 retested plans dropped 35-61% in five days. None went up.
Anyone else seeing similar in their own usage?
The price of intelligence is dropping fast. You can run GPT-4 level models locally today for almost nothing in compute cost. That trend is real and continuing. Just take a look at https://artificialanalysis.ai/models/gpt-4 https://artificialanalysis.ai/models/gemma-4-26b-a4b
And it does look undeniable that LLMs are genuinely useful. The "are they useful at all" question feels settled.
Where I agree is on the financial side — what the top labs are doing does look like a bubble. They're racing toward being first to AGI as if that gives them world domination. That seems delusional, and they'll need to stop or collapse. But collapse won't kill AI or LLMs. It's like saying the dot-com bubble collapse should have killed the internet because it wasn't useful. The internet was useful, dot-com bubble or not. LLMs are useful, bubble or not.
What I expect is things will get less crazy over the next few years, with or without a crash. GPT-4+ and Opus 4+ models have reached genuinely useful levels for knowledge work. And if the trend continues — from GPT-4 expensive in the cloud to Gemma 4 running locally and being smarter — we'll get GPT-5/Opus 4 level models running locally by 2028, and it will keep getting better. At that price point and with local use, AI will be where it should be.
The top labs are burning money like crazy to get there first, as if that gives them a lasting moat. It won't.
What they're trying to do is become the new Google, Apple, or Microsoft. Those companies did achieve sustained moats through speed of execution. I don't see that happening in AI right now.
Though the Anthropic moment this year did make me worry a little...
It's interesting to compare it to electricity. Basically Anthropic was selling a flat fee electricity subscription, and when someone started connecting expensive washing machines (OpenClaw) to their subscriptions, instead of changing the pricing model, they banned washing machines...
I wonder if we will get to "electricity" style pricing for AI. What makes electricity predictable is relatively constant average usage over time + price is manageable. I'm just not buying electrical house heating and manage my electricity spending within some bounds.
With AI the problem is that we are only now getting to useful AI, and for now it's still too expensive to be useful, so they subsidize until they can stabilize at "cheap enough and smart enough" level. But it feels like that's still 2 years away while they are stopping to subsidize now. Will be interesting.
Zettlekast has other benefits for humans though. If your goal is to grasp lot of knowladge oyu need to do it in atomic way, connect mentally to what you already know and do spaced repetition to internalize. Zettlkeaste forces you do it it all as part of organizing. Basically by organizing you make it your own.
Yes AI can help today but it also means it does not stay in your head. Not sure its important if it is in your head or you can call AI at any moment instead of your own memory.
That made me wonder: what is human uptime if measured the same way — against a 24/7 clock?
Agents are more like humans, not SaaS, not only in how to work with them, in other ways too. Does it make sense?
The idea:
Tokens per dollar
Weighted input/output pricing (75/25 assumption)
Benchmark-normalized quality (Arena, Aider, SWE-bench)
Early results surprised me (local often loses economically unless privacy is heavily valued).
I’m mostly looking for critique of the methodology:
Is quality-adjusted tokens per dollar even the right metric?
Is normalizing ELO to % defensible?
What benchmarks am I missing?
Containers are not a big deal when viewed in isolation. But when its common size/standard for all kinds of ships, cranes and trucks, it is a big deal then.
In that sense its more about gathering community around one way to do things.
In theory there are REST APIs and OpenAPI standard, but those were not made for LLMs but code. So you usually need some kind of friendly wrapper(like for candy) on top of REST API.
It really starts to feel like a a big deal when you work in integrating LLMs with tools.