They must have hit really hard scaling limits if the prices were hiked so much so quickly.
They must have hit really hard scaling limits if the prices were hiked so much so quickly.
Can I ask where are you using all those tokens?
I now exclusively use https://omp.sh/ as my harness:
I set it up so it never works in the main branch so subagents etc don't step on each others toes, and only merges back when complete: https://github.com/Daviey/mario/blob/main/.omp/hooks/pre/wor...
A good AGENTS.md is essential: https://github.com/Daviey/mario/blob/main/AGENTS.md
I then provide specifications for what I want, making sure it is unit tested.
For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.
Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.
>They must have hit really hard scaling limits if the prices were hiked so much so quickly.
Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.
The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).
Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.
They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.
They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).
That is shocking. Is it per-token I wonder?
I’m getting 97%.
Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%
[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon
Also why Meta gets a +1, just charge less money on the training path.
If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.
These are not equal.
To be fair, none of us are sure of anything and I think that’s the part that’s most irritating
> ZCode, the GLM coding agent, silently uploads your Git history
And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.