There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token prediction.
There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token prediction.
But would they be where they are if they were not able to borrow heavily from what has come before?
I’m reminded how hard it is to reply to a comment and assume that people will still interpret that in the same context as the existing discussion. Never mind.
If I best you in a 100m sprint people don’t look at our training budgets and say oh well it wasn’t a fair competition you’ve been sponsored by Nike and training for years with specialized equipment and I just took notes and trained on my own and beat you. It’s quite silly in any normal context.
No-one enjoys being taken out context.
But I do accept that given the hostility of replies I didn’t make my point very effectively. In a nutshell, the original comment was that it’s surprising a small team like DeepSeek can compete with OpenAI. Another reply was more succinct than mine: that it’s not surprising since following is a lot easier than doing SOTA work. I’ll add that this is especially true in a field where so much research is being shared.
That doesn’t in itself mean DeepSeek aren’t a very capable bunch since I agree with a better reply that fast following is still hard. But I think most simply took at it as an attack on DeepSeek (and yes, the comment was not very favourable to them and my bias towards original research was evident).