Hm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off.
EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.
I mean, it implies that it has an even better cache hit rate that DS flash, which is impressive, as the chr on DS flash was already really good in my experience
I think the weighted average takes into account all providers (some DS4 flash providers are 'premium' providers and offering higher speeds for higher pricing) and these are tilting the scale
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
Tbh it was also slow because it was being hammered by everyone making use of the free tokens
Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now
Few weeks ago, I wouldn't expect this statement to be true. Accelerate!