Notes on the New Deepseek v3
composio.dev
composio.dev
Also don't just focus on this model but check out what DeepSeek mission is, and the CEO words in the recently released interview. They want to be the DJI / Bambulab of AI, basically: leaders and not followers, and after V3 it's hard to say they don't have the right brains to do that.
Very true. Meta has been disappointing so far, and it takes away from the blog that it starts with a graph of a completely misrepresentative benchmark (MMLU) that shows things like Llama3.1-405b besting Gemini 1.5-pro, 4o-mini above Haiku 3.5, and so on.
But all this means is that the leap in terms of open weights model is even far bigger with Deepseek v3.
- How many 'r's are in Strawberry?
- Finding the fourth word of the response
These tests are at odds with the tokenizer and next-word prediction model. They do not accurately represent an LLM's capabilities. It's akin to asking a blind person to identify colors.
> Here's "strawberry" spelled out one character per line: s t r a w b e r r y
Most LLMs can handle that perfectly. Meaning, they can abstract over tokens into individual characters. Yet, most lack the ability to perform that multi-level inference to count individual 'r's.
From this perspective, I think it's the opposite. Something like the strawberry-tests is a good indicator how far the LLM is able to connect individually easy, but not readily interconnected steps.
> They probably trained the model on a synthetic dataset generated by GPT-4o.
This seems to be the case. I can speculate further. They trained on copyrighted material that OpenAI did not.
It remains to be seen what the pricing will be when run by non-Deepseek providers. They might be loss leading.
The comparison for cheap models should also be Gemini 2.0 Flash Exp. I could see it being even cheaper when it stops being free - if it does at all. There's definitely a scenario where Google just keeps it freeish for a long time with relatively high limits.
these models are being commoditized.
I have a task I use in my work where Gemini 1.5-Pro is SOTA. Handily beating o1, Sonnet-3.5, Gemini-exp and everyone else, very consistently and significantly.
The newer/bigger models are better at reasoning and especially coding, but there's plenty of tasks that have little overlap with those skills.
DeepSeek - 0.14$ per million tokens input, 0.28$ million tokens output (66 tokens per/s)
Fireworks - 0.9$ per million tokens input, 0.9$ million tokens output (23 tokens per/s)
DeepInfra - 1$ per million tokens input, 2$ million tokens output (1.27 tokens per/s)
Compared to Llama 3.1 405B (smaller model than this afaik):
Cheapest is 0.8/0.8$ at 24 t/s all the way to 4$/4$ at 8 t/s
So third party cost seems similar, but there aren't many people hosting DeepSeek right now.
DeepSeek - 0.27$ per million tokens input, 1.10$ million tokens output (66 tokens per/s)
Still much cheaper than the others though for input pricing.
[1] https://api-docs.deepseek.com/news/news1226#-api-pricing-upd...
But still certainly cheaper than everyone else at the moment.
AI slop, I don't trust any of this article, especially the bullets on what made Deepseek "win"
Like with all of these models, we don't know what's in them.