Multiple providers (who need to make a profit) offer the same 4.40 rate for glm-5.2. It's not subsidized.
Deepseek's 0.86 or whatever is likely subsidized but alternate providers offer it for a price comparable to glm-5.2.
Deepseek's 0.86 or whatever is likely subsidized but alternate providers offer it for a price comparable to glm-5.2.
They have published tons of articles dedicated to performance and efficiency engineering. Feel free to have a look...
Does ananyone outside deepseek have a working code for the v4 compressed attention mechanism?
Has any other provider managed to bypass CUDA and program the compute engines in their native assembly language to get 10% more performance out of them?
There is your answer.