I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
I wonder if this was a deliberate move by PRC or really our own fault in falling for the fallacy that more is always better.
And they somehow yolo it for next to nothing?
yes, it seems unlikely they did it exactly they way they're claiming they did. At the very least, they likely spent more than they claim or used existing AI API's in way that's against the terms.
They literally published all their methodology. It's nothing groundbreaking, just western labs seem slow to adopt new research. Mixture of experts, key-value cache compression, multi-token prediction, 2/3 of these weren't invented by DeepSeek. They did invent a new hardware-aware distributed training approach for mixture-of-experts training that helped a lot, but there's nothing super genius about it, western labs just never even tried to adjust their model to fit the hardware available.
It’s also curious why some people are seeing responses where it thinks it is an OpenAI model. I can’t find the post but someone had shared a link to X with that in one of the other HN discussions.
It's extremely cheap, efficient and kicks the ass of the leader of the market, while being under sanctions with AI hardware.
Most of all, can be downloaded for free, can be uncensored, and usable offline.
China is really good at tech, it has beautiful landscapes, etc. It has its own political system, but to be fair, in some way it's all our future.
A bit of a dystopian future, like it was in 1984.
But the tech folks there are really really talented, it's long time that China switched from producing for the Western clients, to direct-sell to the Western clients.
So yes, DeepSeek-R1 appears to be not even be best in class, merely best open source. The only sense in which it is "leading the market" appears to be the sense in which "free stuff leads over proprietary stuff". Which is true and all, but not a groundbreaking technical achievement.
The DeepSeek-R1 distilled models on the other hand might actually be leading at something... but again hard to say it's groundbreaking when it's combining what we know we can do (small models like llama) with what we know we can do (thinking models).
Not that the leaderboard isn't useful, I think "is in the top 10" says a lot more than the exact position in the top 10.
But the claim I'm refuting here is "It's extremely cheap, efficient and kicks the ass of the leader of the market", and I think the leaderboard being topped by a cheap google model is pretty conclusive that that statement is not true. Is competitive with? Sure. Kicks the ass of? No.
Having tested that model in many real world projects it has not once been the best. And going farther it gives atrocious nonsensical output.
Additionally there are claims, such as those by Scale AI CEO Alexandr Wang on CNBC 1/23/2025 time segment below, that DeepSeek has 50,000 H100s that "they can't talk about" due to economic sanctions (implying they likely got by avoiding them somehow when restrictions were looser). His assessment is that they will be more limited moving forward.
OpenAI literally haven't said a thing about how O1 even works.
I'm pointing out that nearly every thread covering Deepseek R1 so far has been like this. Compare to the O1 system card thread: https://news.ycombinator.com/item?id=42330666
Very different standards.
Maybe we don't need momentum right now and we can cut the engines.
Oh, you know how to develop novel systems for training and inference? Well, maybe you can find 4 people who also can do that by breathing through the H.R. drinking straw, and that's what you do now.
Oh dear
There's something wrong with the West's ethos if we think contributing significantly to the progress of humanity is malicious. The West's sickness is our own fault; we should take responsibility for our own disease, look critically to understand its root, and take appropriate cures, even if radical, to resolve our ailments.
Who does this?
The criticism is aimed at the dictatorship and their politics. Not their open source projects. Both things can exist at once. It doesn't make China better in any way. Same goes for their "radical cures" as you call it. I'm sure Uyghurs in China would not give a damn about AI.
Which reminded me of "Whitey On the Moon" [0]
it's quite like Trump's 'CHINA!' yelling
I don't know, just a guess
Do you want an Internet without conspiracy theories?
Where have you been living for the last decades?
/s
https://www.chinalawtranslate.com/en/generative-ai-interim/
In the case of TikTok, ByteDance and the government found ways to force international workers in the US to signing agreements that mirror local laws in mainland China:
https://dailycaller.com/2025/01/14/tiktok-forced-staff-oaths...
I find that degree of control to be dystopian and horrifying but I suppose it has helped their country focus and grow instead of dealing with internal conflict.
The vast majority are completely ignorant of what Socialism with Chinese characteristics mean.
I can't imagine even 5% of the US population knows who Deng Xiaoping was.
The idea there are many parts of the Chinese economy that are more Laissez-faire capitalist than anything we have had in the US in a long time would just not compute for most Americans.
I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.
And since it's a businessperson they're going to make it sound as cute and innocuous as possible
When deciding between mostly like scenarios, it is more likely that the company lied than they found some industry changing magic innovation.
I'm not even saying they did it maliciously, but maybe just to avoid scrutiny on GPUs they aren't technically supposed to have? I'm thinking out loud, not accusing anyone of anything.
Something makes little sense in the accusations here.
https://wccftech.com/nvidia-asks-super-micro-computer-smci-t...
They probably also trained the “copied” models by outsourcing it.
But who cares, it’s free and it works great.
https://wccftech.com/nvidia-asks-super-micro-computer-smci-t...
Chinese guy in a warehouse full of SMCI servers bragging about how he has them...
Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38. "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..."
Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people.
Model parameter count and training set token count are fixed. But other things such as epochs are not.
In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you have.
Also, what if their claim on GPU count is accurate, but they are using better GPUs they aren't supposed to have? For example, they claim 1,000 GPUs for 1 month total. They claim to have H800s, but what if they are using illegal H100s/H200s, B100s, etc? The GPU count could be correct, but their total compute is substantially higher.
It's clearly an incredible model, they absolutely cooked, and I love it. No complaints here. But the likelihood that there are some fudged numbers is not 0%. And I don't even blame them, they are likely forced into this by US exports laws and such.
I don't expect a #180 AUM hedgefund to have as many GPUs than meta, msft or Google.
This is just not true for RL and related algorithms, having more GPU/agents encounters diminishing returns, and is just not the equivalent to letting a single agent go through more steps.
I feel like if that were true, it would mean they're not lying.
Under such dire circumstances, lying isn't entirely out of character for a corporate CEO.
Deepseek obviously trained on OpenAI outputs, which were originally RLHF'd. It may seem that we've got all the human feedback necessary to move forward and now we can infinitely distil + generate new synthetic data from higher parameter models.
I’ve seen this claim but I don’t know how it could work. Is it really possible to train a new foundational model using just the outputs (not even weights) of another model? Is there any research describing that process? Maybe that explains the low (claimed) costs.
Those were probably from OpenAI models. Then they used reinforcement learning to expand the reasoning capabilities.
> You can RL post-train your small LLM (on simple tasks) with only 10 hours of H100s.
https://www.reddit.com/r/singularity/comments/1i99ebp/well_s...
Forgive me if this is inaccurate. I'm rushing around too much this afternoon to dive in.
note: I'm not Chinese, but AGI should be and is a world wide space race.
As OP said, they are lying because of export laws, they aren’t allowed to play with Nvidia GPUs.
However, I support DeepSeek projects, I’m here in the US able to benefit from it. So hopefully they should headquarter in the States if they want US chip sanctions lift off since the company is Chinese based.
But as of now, deepseek takes the lead in LLMs, my goto LLM.
Sam Altman should be worried, seriously, Deepseek is legit better than ChatGPT latest models.
Many "haters" seem to be predicting that there will be model collapse as we run out of data that isn't "slop," but I think they've got it backwards. We're in the flywheel phase now, each SOTA model makes future models better, and others catch up faster.
Just a cursory probing of deepseek yields all kinds of censoring of topics. Isn't it just as likely Chinese sponsors of this have incentivized and sponsored an undercutting of prices so that a more favorable LLM is preferred on the market?
Think about it, this is something they are willing to do with other industries.
And, if LLMs are going to be engineering accelerators as the world believes, then it wouldn't do to have your software assistants be built with a history book they didn't write. Better to dramatically subsidize your own domestic one then undercut your way to dominance.
It just so happens deepseek is the best one, but whichever was the best Chinese sponsored LLM would be the one we're supposed to use.
- OP elides costs of anything at all outside renting GPUs, and they purchased them, paid GPT-4 to generate training data, etc. etc.
- Non-Qwen models they trained are happy to talk about ex. Tiananmen
Since the model is open weights, it's easy to estimate the cost of serving it. If the cost was significantly higher than DeepSeek charges on their API, we'd expect other LLM hosting providers to charge significantly more for DeepSeek (since they aren't subsidised, so need to cover their costs), but that isn't the case.
This isn't possible with OpenAI because we don't know the size or architecture of their models.
Regarding censorship, most of it is done at the API level, not the model level, so running locally (or with another hosting provider) is much less expensive.
Snowden releases?
as DeepSeek wasn't among China's major AI players before the R1 release, having maintained a relatively low profile. In fact, both DeepSeek-V2 and V3 had outperformed many competitors, I've seen some posts about that. However, these achievements received limited mainstream attention prior to their breakthrough release.
Correct me if I'm wrong, but couldn't you take the optimization and tricks for training, inference, etc. from this model and apply to the Big Corps' huge AI data centers and get an even better model?
I'll preface this by saying, better and better models may not actually unlock the economic value they are hoping for. It might be a thing where the last 10% takes 90% of the effort so to speak
I do not quite follow. GPU compute is mostly spent in inference, as training is a one time cost. And these chain of thought style models work by scaling up inference time compute, no?
So proliferation of these types of models would portend in increase in demand for GPUs?
So it's not the end of the world. Look at the efficiency of databases from the mid 1970s to now. We have figured out so many optimizations and efficiencies and better compression and so forth. We are just figuring out what parts of these systems are needed.
They bought them at "you need a lot of these" prices, but now there is the possibility they are going to rent them at "I dont need this so much" rates.
OpenAI will be also be able to serve o3 at a lower cost if Deepseek had some marginal breakthrough OpenAI did not already think of.
This is a net win for nearly everyone.
The world needs more tokens and we are learning that we can create higher quality tokens with fewer resources than before.
Finger pointing is a very short term strategy.
If someone gets something to work with 1k h100s that should have taken 100k h100s, that means the group with the 100k is about to have a much, much better model.
This will be good. Nvidia/OpenAI monopoly is bad for everyone. More competition will be welcome.
they'll be fine: https://www.msn.com/en-us/news/technology/huawei-smic-to-bui...
GPU: nope, that would take much longer, Nvidia/ASML/TSMC is too far ahead
DeepSeek's R1 also blew all the other China LLM teams out of the water, in spite of their larger training budgets and greater hardware resources (e.g. Alibaba). I suspect it's because its creators' background in a trading firm made them more willing to take calculated risks and incorporate all the innovations that made R1 such a success, rather than just copying what other teams are doing with minimal innovation.
I've seen a $5.5M # for training, and commensurate commentary along the lines of what you said, but it elides the cost of the base model AFAICT.
So I doubt that figure includes all the cost of training.
You also need to keep the later generation cards from burning themselves out because they draw so much.
Oh also, depending on when your data centre was built, you may also need them to upgrade their power and cooling capabilities because the new cards draw _so much_.
Claude gave me a good analogy, been struggling for hours: its like only accounting for the gas grill bill when pricing your meals as a restaurant owner
The thing is, that elides a lot, and you could argue it out and theoratically no one would be wrong. But $5.5 million elides so much info as to be silly.
ex. they used 2048 H100 GPUs for 2 months. That's $72 million. And we're still not even approaching the real bill for the infrastructure. And for every success, there's another N that failed, 2 would be an absurdly conservative estimate.
People are reading the # and thinking it says something about American AI lab efficiency, rather, it says something about how fast it is to copy when you can scaffold by training on another model's outputs. That's not a bad thing, or at least, a unique phenomena. That's why it's hard talking about this IMHO
To know that this would work requires insanely deep technical knowledge about state of the art computing, and the top leadership of the PRC does not have that.
https://x.com/sivil_taram/status/1883184784492666947?t=NzFZj...
https://medium.com/the-generator/deepseek-hidden-china-polit...
But also the claimed cost is suspicious. I know people have seen DeepSeek claim in some responses that it is one of the OpenAI models, so I wonder if they somehow trained using the outputs of other models, if that’s even possible (is there such a technique?). Maybe that’s how the claimed cost is so low that it doesn’t make mathematical sense?
In theory I could run this one at home too without giving my data or money to Sam Altman.
also deepseek is open-weights. there is nothing preventing you from doing a finetune that removes the censorship. they did that with llama2 back in the day.
This is an outrageous claim with no evidence, as if there was any equivalence between government enforced propaganda and anything else. Look at the system prompts for DeepSeek and it’s even more clear.
Also: fine tuning is not relevant when what is deployed at scale brainwashes the masses through false and misleading responses.
The enforcers identity is much more important.
If you think these tech companies are censoring all of this “just because” and instead of being completely torched by the media, and government who’ll use it as an excuse to take control of AI, then you’re sadly lying to yourself.
Think about it for a moment, why did Trump (and im not a trump supporter) re-appeal Biden’s AI Executive Order 2023 ? , what was in it ? , it is literally a propaganda enforcement article, written in sweet sounding, well meaning words.
It’s ok, no country is angel, even the american founding fathers would except americans to be critical of its government during moments, there’s no need for thinking that America = Good and China = Bad. We do have a ton of censorship in the “free world” too and it is government enforced, or else you wouldnt have seen so many platforms turn the tables on moderation, the moment trump got elected, the blessing for censorship directly comes from government.
What do you think they will do with the AI that worries you? They already had access to Llama, and they could pay for access to the closed source AIs. It really wouldn't be that hard to pay for and use what's commercially available as well, even if there is embargo or whatever, for digital goods and services that can easily be bypassed