Any idea on how they get the pricing so efficient? Their artificialanalysis graph has them on par with GLM 5.3 but at less than 10% of the price despite being larger than GLM 5.3.
It sounds like they are pricing very well because OpenAI and Anthropic have been scamming us for years. Once the infrastructure is in place, electricity cost is the only concern. China supports businesses and gives them a lot of incentives to lower their costs. And who knows what's being provided to them without anyone knowing. All in the name of winning the race.
If that is true, why are the open model inference providers on OpenRouter (and on their own website) so expensive then? It must be the hardware. I've checked and a lot of it it just the HBM, with the NVidia tax being a smaller but also large factor.
Nvidia's top AI chip, Rubin, sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro could generate roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price. That's a payback of a few weeks in theory.
So yes, we are getting scammed by American SOTA.
Interesting. Then my theory is that the free market inference providers have low utilisation (e.g. 20% avg; 100% at peak), slowing their amortization rate. The OpenAI/Anthopic rates must then have been based on a time when there utilisation was peaky and mostly idle (e.g. 10/20%). That would maybe explain the 10-20x cost difference between coding plans and API.