Most people either don’t have the money to buy hardware to run open models and/or don’t have the utilisation to make renting cost effective.
Most people either don’t have the money to buy hardware to run open models and/or don’t have the utilisation to make renting cost effective.
you can use these in hermes, cursor, openclaw, opencode, etc with 2 lines of config that claude code will happily do for you if you ask
GLM 5.2, deepseek 4 Flash and the newly released Hy3 are Opus 4.8 and Sonnet 4.6 level models at a tiny fraction of the cost.
I'm on the $200 claude plan and blew through my weekly limits with Fable in a day, then ended up wasting $20 with opus 4.8 overages in an hour to finish out work in active sessions. Since then I've been using GLM 5.2 with openrouter + opencode and am spending less than $5/day for equivalent output.
But it's a very competitively priced model other providers can offer (since it's open) so it's a much cheaper alternative than claude in practice.
I assume that is what they meant.
You'd need at least 24 new M4 Max Studios, or 16 new M3 Ultra Studios, or 3 used 512 GB M3 Ultra Studios just to power one GLM 5.2 instance. And even then, you're probably looking at < 5 tokens per second.
Personally, I think it makes way more sense to pay a model provider $3/1M tokens.
You'll never get proper price competitive utilization on personal hardware vs a cloud inference provider that can batch and pipeline requests optimally to maximize utilization, unless you yourself start running batch jobs.
Even once local hardware and models catch up to todays frontiers, by that time there will be 10x better cluster sized models available at a similar discount.