Ox-Alpha Is GLM?
dejan.ai
dejan.ai
https://files.catbox.moe/k52n6k.png
The upper chart shows the availability of Ox Alpha and the lower chart shows the availability of GLM 5.3 by Z.ai. They had a blip at exactly the same time.
My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot.
MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant than the lab behind ox-alpha, so less likely).
The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
They have had vision models before just not their flagships
It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.
https://www.tomshardware.com/tech-industry/artificial-intell...
I believe it's GLM 5.3 Flash or Air.
- The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely
- The model is extremely efficient, beyond anything we've seen so far
- Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)
> 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day.
https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300
But while we’re “guessing”: Xiaomi MiMO
Also Claude and the Chinese models really like to say "Let me X", Geminis preference is "I will X".
They will never ever do like this.
They for sure would add architecture ideas from other research/models but thats it.
They were just trolling.
Ox Alpha
"The reasoning model that appeared out of nowhere. Built for code, long-horizon agents, and a million tokens of context. Nobody knows who made it — everyone wants to try it."
Its amazing to me that providers haven't added any sort of masking of the prompt in the thinking traces to avoid prompt extraction via this sort of trivial attack
I wish projects would check before stealing a name.
But I don't know. There are vagueposts on X about this model running on two DGX Sparks. If they did the same scale-up with Air instead of Flash, that would probably be about right.
We're lucky if it'd fit in 1 DGX Spark. Laptops - nah, unless you mean like an M5 Max with 128GB of RAM then maybe.
> Thats the reason behind the hype.
The hype is imagine DeepSeek Flash before the price increase with even better performance. It'd be like unlimited Sonnet.
I think something also went wrong with the Ox provider last night (at least on OpenRouter), for a few hours it wouldn't accept tools. Zero change to the harness while I slept and it was back working again the next morning.