2,886 karma · joined November 27, 2010
[1] https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...
Last year, Anthropic cut off Windsurf the same way [2]. Fortunately, unlike Anthropic, OpenAI allows you to use your subscription with other harnesses, including Cursor.
Of course today, Anthropic is taking the "high road" and is OK with with SpaceX [3] including their models in Cursor, because they need SpaceX's compute.
[1] https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...
[2] https://techcrunch.com/2025/06/03/windsurf-says-anthropic-is...
When GPT 6 comes out, would you expect the top thread to link to OpenRouter?
- https://api-docs.deepseek.com/
- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)
But even in this very post, you can see that Max was actually cheaper than High.
If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.
For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more.
EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All charge exactly the same as the first-party. Does that mean that their costs are the same? Seems quite unlikely. It's simply not an efficient market, yet.
Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.
Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.
Remember they ~doubled the price going from GLM 5 to GLM 5.2, despite same [1] cost of inference.
[1] GLM 5.2 is actually slightly more efficient, thanks to baked in indexer cache.
DeepSeek V3.2 which uses DSA only (sparse attention, but without compression from HCA and CSA) is a smaller model but uses 10x more memory at 1M context window compared to DS V4 Pro.
Also, I have to say, DeepSeek's API has a very good cache hit rate. With the same workload, I see ~80% KV cache hit rate with the DS API vs ~50% with the major western inference providers for open weight models.
Inference stack efficiency: Many of these providers take off the shelf sglang / vllm / trtllm and hope for the best. Meanwhile DeepSeek team is known for pushing the boundary of optimizations.
Now, sglang and vllm are great pieces of software, but take DeepSeek's Sparse Attention (DSA). Introduced 1.5 years ago (https://arxiv.org/abs/2512.02556), used by DeepSeek 3.2, GLM 5, DeepSeek V4. Only now is it slowly strating to get optimized in the major inference engines: (https://github.com/sgl-project/sglang/issues/19380 https://github.com/sgl-project/sglang/pull/22851 etc.). Of course, DS V4 adds extra optimizations into the model architecture on top of DSA, and those will take more time to be taken full advantage of by the open source inference engines.
Privacy: Betting that people will pay extra for inference hosted outside China. This is especially true with DeepSeek, because DeepSeek is transparent about using API data for model improvements.
And few other things (scale (matters a lot for MoEs), reliability, soft enterprise lock in, etc.)
---
There is also, likely, tacit collusion at play here. Look at GLM 5 and GLM 5.1 prices. GLM 5 and 5.1 cost the same to run, but providers decided to charge much more for 5.1 because it is much better model, and because Z.AI raised their price as well.