This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API prices and speeds mean nothing for open source models.
This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API prices and speeds mean nothing for open source models.
(there's a table which shows comparison between vendors)
Also, it seems there's a general one as well (for all kimi models?): https://github.com/MoonshotAI/Kimi-Vendor-Verifier
We use three classes of signals:
* Tool-calling success and reliability from real traffic
* Provider performance metrics such as throughput and latency
* Benchmark and evaluation data as it becomes availableLooking at openrouter [1], some of the cheaper offerings are for quantized models. Not sure how much intelligence is lost in quantization. And they are not 3 times cheaper. Where did you find 3x lower prices for APIs? I am considering skipping open router and using them directly for that price.
edit:
I see, croft [2] 8bit for $0.50/$0.08/$2.20
I do not have GLM 5.2 numbers because the whole default max setting is overkill. But GLM 5.1 numbers had it at 12x cheaper then API rates. And about 2.5x more tokens vs zai their own subscription service.
Yes, its FP8 but lets be honest, do we know for sure that even zai runs at FP16? I learned a long time ago with Claude and Codex how much cheating happens on model levels, even from the big boys.
On top of that, the cloud offering doesn't seem that well-run, they randomly blocked a colleague's API key for a couple days without any heads up, had a weird rate limiting bug and they have been deprecating models without redirects with very short notice, all while taking weeks to onboard new models. I assume some of these problems would be addressed if we had an SLA/enterprise contract.
It's a promising idea though. They offer a $5 trial credit (with an aggressive rate limit) though so no harm in trying it out.
Absolute false information.
From my usage panel for this month:
* Total Tokens 1.1B * Cached Tokens 1.0B 97% of prompt tokens * Cost energy pricing $26.58
The energy pricing is higher then what i actually pay because its a mix of token billing and partial subscription (60% extra "power").
From the $50 subscription, i have about 3/4 left (4.21 of 16.0 kWh used this billing cycle). Used $5.5 in token billing.
That was running 82.0% GLM 5.1, and 18% GLM 5.2. Yes, i have been busy ;)
My actual usage if we look in dollar value was ~ $18.
For your information, that is cheaper the MiMo v2.5 Pro from Xiaomi as there i was doing around 450.000t per cent. And they have the same 75% cheaper prices like DeepSeek. MiMo has a issue with cache retention between session prompts what hurts them vs DeepSeek. Yes, DeepSeek v4 Pro is 2.5x cheaper but nowhere near GLM 5.1, and especially not GLM 5.2.
In case your wondering, zai subscription light is about 80m token / week limit. So on a token/cent price, neutralwatt is about 3x cheaper (and not 5h, week limits to maximize/frustrate).
> all while taking weeks to onboard new models.
Took them 1 day to include GLM 5.2 ... Yes, the remove old models fast because they do not have the server capacity to keep old models around.
> I assume some of these problems would be addressed if we had an SLA/enterprise contract.
Its a small team, not a big huge company. From my experience so far, seen a 2 timeouts, and sometimes slow speeds as servers get overloaded. For what i am paying for GLM ~5.1~ 5.2 ...
I am not sure why the small team argument is relevant. This is a crowded market, there are dozens if hundreds of third party inference providers in the world right now. I'm glad that's a good excuse that works on you but I'm not sure why the average user should care.
Then you actually use the service and see how much tokens you use on average. You calculate the token use vs what you pay. And this gives you a stable number to compare different services and model with, if you want the token cost. This is basic school level reasoning and calculation.
> I am not sure why the small team argument is relevant.
This is relevant to the previous poster his question regarding support and SLA/enterprise support.
> Your reply doesn't seem to be in good faith.... I'm glad that's a good excuse that works on you ...
Question: Do you have a issue with communicating with other people in real life?
Claude Shannon is rolling in his grave.
https://en.wikipedia.org/wiki/Rate%E2%80%93distortion_theory
Shannon would be pointing out that if you can throw away half the model without apparent degradation, we're nowhere near packing in all the information we could in training. There must be a better arrangement than we've currently got.
So you could end up paying more for unquantised weights, only to get silently hit with a quantised KV cache...
I've tried a number of these, and the learning curve is very steep compared to "install Claude Code and pay $100/mo". There is no way saving me $50/month matters compared to figuring that out.
https://docs.z.ai/devpack/tool/claude
Here's my setup. I add this to my .bashrc
export ZAI_API_KEY="your_key_here"
alias claudez='ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY" ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic" ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5.2[1m]" ANTHROPIC_DEFAULT_SONNET_MODEL="glm-4.7" ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-4.7" claude'
Then I just run claudez
pro tip the same thing works with deepseek https://api-docs.deepseek.com/guides/anthropic_api
Even more pro tip: Claude Code can set this up for you haha
Unless this were a massive differentiator, people aren't going to be "talking about it" the way GP suggests!
It's crazy that apparently writing software without knowing how to edit a single config file is normal now.
At 40, I could easily configure claude code to use another model, even if there weren't any official guides with a bit of MITM fun, but I don't want to invest my attention / heavily use something that will most likely break in the near future.
I'd argue math is significantly more complex than language.
1) You haven't even heard of it.
2) You have to know to look for both GLM and Z.ai. These are usually in the same article when reporting about GLM is written, at least.
3) You have to understand there could be a benefit in trying it; you have to want to try it for some reason. Their own blog post puts it below Opus 4.8 in each of the three benchmarks they used.
4) You have to figure out the pricing, which isn't obviously in the blog post...
5) When I first went to Z.ai, I got an error popup (not logged in): "You do not have permission to access this resource. Please contact your administrator for assistance." I am using a personal computer...
6) When I typed something in the resultant field and pressed enter, I got "Clear Current Chat? To start a new chat, your current conversation will be discarded. Sign in to save chats"
I think today's article helped with 1 and 2, which helps their top of funnel. But they're fighting a big uphill battle.
There's ZCode (https://zcode.z.ai). Which is like the Codex App.
That's as "easy" as it is for non-devs that you're complaining about.
Yes, there is. It's called Claude Code. Point it at the HuggingFace URL and say "Download these weights and build whatever is needed to run them, then test the model."
(In all seriousness, I agree this is a problem. That capability is too powerful not to take advantage of, though. Nobody needs to struggle with this sort of thing anymore, but yes, obviously, it should happen in a VM or at least a container.)
I'd pay for an out of the box solution. i.e. an Installer with updates
Wasn't this released like 2 days ago? Everyone is still evaluating and playing around with it, things like the submission is just starting to come out. Give it some days at least before jumping to conclusions, ideally weeks.
Now, maybe GLM 5.2 is close to Opus 4.7, but I don't wanna keep checking them and keep finding that they're still benchmaxing and aren't at GPT (my choice) or Opus level. The boy who cried wolf, I guess.
1. Keeping your data private on in the US
2. Not training on it
3. Not quantizing the model
4. Offer reasonable latency adds rate limits
With that said, I'm excited to try GLM 5.2 because I still end up reaching for Opus and GPT 5.5 for many tasks because the open models tend to get stuck more often on complex problems.
Do note that GLM is not multi modal, which can be a deal breaker. And these open models are not good outside coding.
I wish I had the time to set it up and work on side projects but unfortunately life and work have been crazy (as I'm sure many here feel). That's why I asked for anecdotes about it.
https://github.com/QuantiusBenignus/Zshelf/discussions/2
Not accounting for hardware, of course :)
Not accounting hardware in my costs, since I didn’t buy my hardware for running models. Running models is just something it can do in addition to what I got it for.
Nvidia GPUs are much more efficient than Apple hardware for inference(and training).
link?
> Why
imho everything but opus produces unusable code (fable was even better...), eg gpt5.5 seems to write the absolute worst code that still technically solves the problem; tbh I'd be totally willing to trade "raw intelligence" for "code taste"
more labs need to figure out whatever anthropic did to destroy everybody else on frontiercode bench
regarding edge cases -- less is more in my experience, as removing is harder than adding