But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
$500 for the old $200 is definitely a fumble.
Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.
There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.
The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.
RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.
I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.
> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
Now you can use it in coding harnesses that call the API.
Why? You can use in codex, right?