GLM-5.3-Flash Intelligence, Performance and Price Analysis
artificialanalysis.ai
artificialanalysis.ai
- costs per task $0.05 vs $0.09
- speed 130 vs 88
- where GLM has only 5 more intelligence point: at this point few point is meaningless for most of models
https://artificialanalysis.ai/models/comparisons/glm-5-3-fla...
Been using Luna exclusively since the price drop, and i've been very satified with all tasks from planning, writing code, and other agent tasks. (just change thinking level from low <-> ultra)
---
btw, I did try out Ox Alpha, the coding feels good but still not way better for me to switch to it.
I have found that sometimes a smaller model with max reasoning is actually more expensive than using the next tier model with a lower reasoning effort. It’s certainly faster.
So Luna is competitive because a few weeks ago they did a 80% price drop?
Many here said that 80% drop was not a move against Anthropic but a move against chinese models and your comments indicate that's the case.
Probably shouldn’t say this here but I’ve been planning to up my $20/mo exploratory ChatGPT subscription to the $100/mo tier as soon as I hit my cap. Between the progress and quality of Luna and their continuous resets, it’s been a few months now that I’ve lived off the $20 tier, frankly waiting for the need to upgrade, credit card in hand.
I’m always trying new models, like many of us here, but the price is just so good for a well balanced, American, hosted model.
For me at this point, most of newer models are capable enough, I focus more on $ and how much I can save.
I signed up for their Lite plan when it was only $28 for the whole year (less than $3/mo). Definitely very happy with that purchase!
In general I really don't mind waiting 5, 10, 40 minutes. There's other things I can look at, other plans or assessments or outputs aplenty stacking up. Its baffling beyond words to me that anyone would take speed over good output. Surely the better output is going to save enormous time in the long run, have better outcomes. What is it that addicts people so much to speed, especially when the difference is between fast and very fast?
You might think faster is better when you vibe code or monitoring. But with background agents, the faster the speed, the more jobs it can perform.
The speed won't matter as much for your personal projects, but if you want to handle enterprise level request, the faster the better.
Say you queued up all messages for a task in say Kafka, you got workers calling AI agents. You will have thousands of messages to do, and the faster the AI agents can do its work the better you can clear the queue.
Intel Cost
lig per
ence Task Speed
Gemini 3.7 flash 56 0.40 338
GLM 5.3 flash 57 0.09 49
Factor 1 4.4 6.9
I have both GLM & Gemini in a subscription and see no reason for choosing GLM 5.3 Flash. Working with de speed of Gemini 3.7 Flash is such a delight that I accept the hassle of working with Antigravity CLI, coming from Claude Code which I use for GLM.GLM-5.3-Flash takes 7x (relative to Gemini and Sol) per task. So, it's cheaper, if you don't value your time! Don't value real-time workflows, don't value iteration speed, etc. So, doesn't seem very suitable for interactive or agentic work to me.
But having an ultra cheap model for async stuff is always very nice. (Still, the last few weeks feel less about tech and more like a contest between who can afford to give the biggest discounts!)
--
I also like DeepSwe[2], although they measure Output Tokens and Agent Steps, which are misleading when one model has a much faster output speed. (e.g. on their metrics Gemini looks slower, because they don't account for that.)
[0] Time per Task - https://artificialanalysis.ai/?models=glm-5-3-flash%2Cgemini...
[1] Output Tokens Per Task - https://artificialanalysis.ai/?models=glm-5-3-flash%2Cgemini...
Those are Musk-like businesses, on steroids.
Not even Tesla has been profitable compared to the capital raised and the debt issued.
OpenAI and Anthropic are already in a ~200bil hole from previous model iterations and are committing to trillions of additional spending
OpenAI spent more TBPN than kimi spent on training K3
They are by all accounts, not. Z.ai for instance is a public company according to wikipedia. Moonshot AI is private but all their investors are private companies. Alibaba, as we all know, is a massive publicly traded tech conglomerate.
Moreover even if we take the more charitable view that they're controlled by the CCP, and therefore will continue releasing models for free, that seems as questionable as the prospect that private investors will continue shoveling money into anthropic/openai.
Don't forget, Jack Ma of "publicly owned" Alibaba, had to go into classroom time out after seemingly forgetting that its classroom capitalism and not real world capitalism.
2. people tend to ignore this, but the salary budget of a US frontier lab and chinese frontier lab is nowhere comparable, the first can easily outdone the later by 100x.
3. us labs, like other US style startups, always throw ton of money to capture the market. I don't see the chinese company doing the same scheme at all.
so, surely chinese AI providers also lost money making new models, but they are not spending nearly as much as US ones.
>2. people tend to ignore this, but the salary budget of a US frontier lab and chinese frontier lab is nowhere comparable, the first can easily outdone the later by 100x.
Both arguments make it seem like there's a double standard for american vs chinese AI companies, where american labs are held up to strict standards for profitability, but chinese labs get a pass because [insert handwaving about how some aspect of chinese labs is different]. Let's do apples to apples comparisons here, what are both sides' run rates and revenue growth prospects?
>3. us labs, like other US style startups, always throw ton of money to capture the market. I don't see the chinese company doing the same scheme at all.
Right, instead they're releasing their models for free so competitors can undercut them on inference. American labs' prospect of "there are open models 90% as good but cost less" might seem bad, but chinese labs' prospect of "there are companies offering the exact same models but aren't on the hook for r&d spend" seems even worse.
Not really, in a way. Things just cost far more in the US than in China; has pretty much always been the case, far back as I can recall. The Chinese state heavily invests in anything it wants to succeed at, and it has the resources to throw. Cost of living is generally wildly lower in China, along with salaries (although it's also pretty location- and role-dependent).
Overall I'd say labs are far cheaper to run in China than in the US, in more than just from the finances angle.
PRC AI have lower opex and capex, i.e. export controls means they couldn't be trillions in the hole on inflated hardware in the first place. They only need to extract a few 10s of billions from domestic market have a healthy runway. If investors/gov wants to throw in a few billion to treat as utility, whatever, it's still rounding error.
This seems like a double standard given that china is still building coal.
It's really simple: if they truly get to human-level AI (or even superhuman AI), then money and debts no longer matter, since our current economic system will be obsolete. They are betting everything on this outcome.
I don't know if they will manage to do it before their debts have to be repaid, but considering the rate of acceleration in the past few months, there is a non-trivial chance that they will, IMHO. We will see.
LLM has nothing to do with AGI.
We do know, that the ceiling is below AGI. And it's not a matter of opinion - LLMs can not achieve AGI due to their design. Anyone telling you otherwise is lying to you.
And it doesn't matter how many or how severe bugs they can find, because it's not about what they produce, but how they produce it.
[citation definitely needed]
> And it doesn't matter how many or how severe bugs they can find, because it's not about what they produce, but how they produce it.
AGI is defined by the practical outcomes, not by the way the outcomes are achieved. You have no way to know that scaled-up Transformers predicting the next token will never result in human-level intelligence, since we currently have no idea where the ceiling of that approach is.
If you're not even familiar with how LLMs work, perhaps you should restrain yourself from confidently talking about this topic until you educate yourself. You're only spreading misinformation.
> AGI is defined by the practical outcomes, not by the way the outcomes are achieved
That for sure would be a very convenient definition, especially for all those AI labs trying to convince investors that they achieved AGI. Unfortunately everyone knows, that knowing the right answer isn't the same as knowing where that answer came from.
I am very familiar with how LLMs work, and I am telling you that there is no consensus that they cannot achieve AGI in the machine learning community. Some people think so (such as Yann LeCun), others disagree. We just don't know yet.
> That for sure would be a very convenient definition, especially for all those AI labs trying to convince investors that they achieved AGI. Unfortunately everyone knows, that knowing the right answer isn't the same as knowing where that answer came from.
AGI is defined by capabilities, not methods.
And the model isn't even shown in the speed bar chart just below. Such slop (the artificial intelligence website linked)