Their parent hedge fund company isn't huge either, just 160 employees and $7b AUM according to Wikipedia. If that was a US hedge fund it would be the #180 largest in terms of AUM, so not small but nothing crazy either
Their parent hedge fund company isn't huge either, just 160 employees and $7b AUM according to Wikipedia. If that was a US hedge fund it would be the #180 largest in terms of AUM, so not small but nothing crazy either
The negative downsides begin at "dystopia worse than 1984 ever imagined" and get worse from there
https://x.com/angelusm0rt1s/status/1881364598143737880
Be careful
Oh please, current and next gen LLMs will be absolutely fantastic for education:
https://x.com/emollick/status/1879633485004165375
Personalized tutors for everyone.
It's indeed very dystopia.
While it is hard to predict the future, a good bet is that global trade will win out in the end on a long time frame.
Arguably China doesn't have the technology required to manufacture 30-series GPUs with the yield or unit cost Nvidia did. I wouldn't hold my breath for Chinese silicon to outperform Nvidia's 40 or 50 series cards any time soon.
Both R1 and V3 say that they are ChatGPT from OpenAI
If you have a model that can learn as you go, then the concept of accuracy on a static benchmark would become meaningless, since a perfect continual learning model would memorize all the answers within a few passes and always achieve a 100% score on every question. The only relevant metrics would be sample efficiency and time to convergence. i.e. how quickly does the system learn?
You say it as if it's an easy thing to do. These things take time man.
I personally would have gone for search/reasoning as has been done. It's the reason path.
DeepSeek is a Chinese AI company and we're talking about military technology. The next world war will be fought by AI, so the Chinese government won't leave China's AI development to chance. The might of the entire Chinese government is backing DeepSeek.
There's a lot more to making foundation models and Deepseek are very much punching well above their weight
The key insight is that those building foundational models and original research are always first, and then models like DeepSeek always appear 6 to 12 months later. This latest move towards reasoning models is a perfect example.
Or perhaps DeepSeek is also doing all their own original research and it’s just coincidence they end up with something similar yet always a little bit behind.
But Google, OpenAI and Meta have chosen to let their teams mostly publish their innovations, because they've decided either to be terribly altruistic or that there's a financial benefit in their researchers getting timely credit for their science.
But that means then that anyone with access can read and adapt. They give up the moat for notariety.
And it's a fine comparison to look at how others have leapfrogged. Anthropic is similarly young—just 3 and a bit years old—but no one is accusing them of riding other companies' coat tails in the success of their current frontier models.
A final note that may not need saying is: it's also very difficult to make big tech small while maintaining capabilities. The engineering work they've done is impressive and a credit to the inginuity of their staff.
We all benefit from Libgen training, and generally copyright laws do not forbid reading copyrighted content, but to create derivative works, but in that case, at which point a work is derivative and at which point it is not ?
On the paper all works is derivative from something else, even the copyrighted ones.
There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token prediction.
But would they be where they are if they were not able to borrow heavily from what has come before?
I’m reminded how hard it is to reply to a comment and assume that people will still interpret that in the same context as the existing discussion. Never mind.
If I best you in a 100m sprint people don’t look at our training budgets and say oh well it wasn’t a fair competition you’ve been sponsored by Nike and training for years with specialized equipment and I just took notes and trained on my own and beat you. It’s quite silly in any normal context.
No-one enjoys being taken out context.
But I do accept that given the hostility of replies I didn’t make my point very effectively. In a nutshell, the original comment was that it’s surprising a small team like DeepSeek can compete with OpenAI. Another reply was more succinct than mine: that it’s not surprising since following is a lot easier than doing SOTA work. I’ll add that this is especially true in a field where so much research is being shared.
That doesn’t in itself mean DeepSeek aren’t a very capable bunch since I agree with a better reply that fast following is still hard. But I think most simply took at it as an attack on DeepSeek (and yes, the comment was not very favourable to them and my bias towards original research was evident).
This is one message of the founders of Mistral when they accidentally leaked one work-in-progress version that was a fine-tune of LLaMA, and there are few hints for that.
Like:
> What is the architectural difference between Mistral and Llama? HF Mistral seems the same as Llama except for sliding window attention.
So even their “trained from scratch” models like 7B aren’t that impressive if they just pick the dataset and tweak a few parameter.
And I have no idea what you mean by "they just pick the dataset". The LLaMA training set is not publicly available - it's open weights, not open source (i.e. not reproducible).
https://epoch.ai/gradient-updates/how-has-deepseek-improved-...