DeepSeek could represent Nvidia CEO Jensen Huang's worst nightmare
marketwatch.com
marketwatch.com
The demand for Nvidia GPUs should now go up. Now anyone can run a GPT-like model by themselves. It's a prime time for businesses to start investing and setting up on-prem infra for that. I know some have been avoiding ChatGPT due to legal concerns and sensitive data.
this conversation reminds me of people when the PS2 came out saying that by 2010 games would look literally better than real life, because they thought graphics quality would exponentially improve...
What was amazing output last year is today's slop.
NV is in no immediate danger; medium/long term, to anyone knowing what "75% gross margin" means it's no secret that there will be serious threats. You don't need DeepSeek for that particular realization.
History is littered with things that were made orders of magnitude better, people who claimed nobody would be happy until that happened, and then nobody writ large cared .
Tech in particular has a grand history of claiming we need to produce the technically best thing, but in practice, the market almost always chooses somewhere in the range of useful products, which is not the same at all - the technically best thing has rarely, if ever, become the market leader.
If there was serious competition for cards to do GPU training that might spell trouble for Nvidia. But so far we only went from needing an absurdly huge cluster of Nvidia cards to a huge cluster of Nvidia cards, with Nvidia having a near-monopoly on cards used in deep-learning training all the way down to networks that fit on single cards.
If the training was too expensive previously is it now cheap enough or does it need to be even cheaper? As if it weren't too expensive everyone who were going to invest did. And if it is still too expensive, well prices have to go lower or more efficiency has to be found.
I am not entirely sure if there is huge amount of unmet demand for training.
It's an objective fact that many people find these things useful. Whether or not they're willing to PAY is another issue.
That day is certainly coming, but today is not that day
The fact that DeepSeek-R1 is so much better than DeepSeek-V3 at various important tasks means that Chain-of-though / thinking-before-answering models are better. But they are also more compute intensive at inference time than their instruction non-thinking counterparts.
So even if the DeepSeek-V3 pretraining + GRPO COT post-training procedure was cheaper than anticipated to reach o1 grade performance, inference is still costly, even if you use a distilled model.
Also, inference is easier to get disrupted in just because of how it works.
All told, this makes it a much less likely play.
Finally, the assumption that it goes up assumes people want on-prem infrastructure for this (in your case). Maybe true, maybe not.
Overall, going from being a "sure" thing selling 200m worth of clusters to random companies hand over fist to not necessarily being able to do that is definitely "bad'.
There are other common claims about why it will increase demand, and they rely on different assumptions (that you won't hit a good enough point quickly, etc)
Or will all of the efficiency savings get immediately absorbed by the demand for better performance and feed the demand for inference?
But i agree in practice that over time nvidia's path now depends heavily on the answer to "Or will all of the efficiency savings get immediately absorbed by the demand for better performance and feed the demand for inference?"
Before, their next 5-10 years did not really depend on meaningful efficiency savings existing and getting them absorbed - nobody expected meaningful efficiency savings.
Now it does depend on that
Those savings don't usually make national news because they aren't the target. They aren't "man on the moon" moments, they are "our cannonball flies 96% as far as yours using 30% less powder" moments.
This is still a race for tools that essentially brute force their way to greater utility. Until we hit a utility quotient or hit a wall, the race doesn't really end. The quickest and easiest path to advantage in the race is still more compute.
I'm aware - i've run the relevant teams at large scale companies like Google :)
Is your claim anyone has made 98% total vs, say, 5 years ago?
If so, evidence please?
Otherwise, seems sort of irrelevant.
AFAIK, nobody has come anywhere close to this, and the only possible way they would is through externalizing roughly all costs onto someone else.
(IE certainly google TPU's are much more efficient than 5 years ago, but since you have to account for development cost, they are not saving 98% vs 5 years ago)
Nvidia will still sell many chips. It just won’t be the only one capable of selling them. The hype is gone. The moat is gone. The excitement and enthusiasm from investors is gone. This is called a correction.
CUDA begs to differ.
Na, we have a long way to go with models, especially when you start adding different modes. We'll still need a metric shitload of compute for a long time.
But there's a more subtle point here which I don't see a lot of people talking about, maybe because they know more about this than me. Why wouldn't frontier model developers take DeepSeek R1 techniques and make their models even better or even larger? Or another way to ask: Are the DeepSeek R1 innovations only for making models cheaper (and slightly worse) or can the algorithms developed by the DeepSeek team be scaled up to make more powerful frontier models? Leading edge model developers don't just care about cost, they also want to release models with maximum capabilities. And as we've seen over the past few years, they're willing to pay almost anything to achieve this.
I think most AI researchers know there are still many things to explore in this space so disruptive innovations shouldn't be seen as a bubble burst but as more opportunity.
NVIDIA is in an extremely strong position right now. Even if someone has a major breakthrough on the hardware design side that dramatically lowers the cost of compute for AI workloads (which is highly unlikely), NVIDIA will just create their own implementation that will outperform the original since they have a stranglehold on an entire stack-up of technology: circuits, drivers, libraries and software.
I don’t think the Nvidia CEO is losing any sleep. He doesn’t have a planet-sized ego, like some. They are still selling pickaxes during a gold rush, as the saying goes. And the stock is still up 100% in the past year. CEOs don’t control the stock price (much as some try), anyway, the market does. Also, the market is insane. You also can’t control the competition, or what new innovations come along. You just have to run a good business, which they are doing.
Innovation is good for all players. Reality-checks are good. Nothing here is unexpected. There are lots of smart researchers in the world. Software innovations that make better use of hardware are expected. People just like drama.
Like if Tim does a better job at a skateboard trick that has always been Tom’s thing, people want to be that kid who is the one to say, “Oh, snap!!” and won’t stop talking about it at school, because they were there. And how it’s so mind-blowing and previously inconceivable.
I understand if new CPU/GPU can outperform Nvidia and DeepSeek was developed using another GPU, but this is not the case. Lower requirement for higher performance historically never reduced the need for computational capacity. There is very poor reasoning for this move other than purely “technical” (trading-wise) reasons.
Their paper (arxiv 2412:1947) explains they used 2048 H800s. A computer cluster based on 2048 GPUs would have cost around $400M about two years ago when they built it. (Give or take, feel free to post corrections.)
The point is they got it done cheaper than OpenAI/Google/Meta/... etc.
But not cheaply.
I believe the markets are overreacting. Time to buy (tinfa).
If that were the case, Cooler Master would be the trillion dollar company, lmao.
Nothing prevents Cooler Master from releasing a line of GPUs equally performant and, while at it, even cheaper. But when we measure reality, after the wave function of fentanyl and good intentions collapses ... oh yeah, turns out only nVidia is making those chips, whoops ...
[0] https://www.tomshardware.com/news/price-of-nvidia-compute-gp...
They estimated $200k for a single NVIDIA GPU-based CPU complete with RAM and networking. That's where my number came from. (RAM and especially very-high-speed networking is very expensive at these scales.)
"Add it all up, and the average selling price of an Nvidia GPU accelerated system, no matter where it came from, was just under $180,000, the average server SXM-style, NVLink-capable GPU sold for just over $19,000 (assuming the GPUs represented around 85 percent of the cost of the machine)"
That implies they assumed an 8-GPU system. (8 × $19,000 = $152,000 ≈ 85% × $180,000)
Yea, a node in a cluster costs as much as an American house. Maybe not on its own, but to make it useful for large scale training, even under the new math of deepseek, it costs as much as a house.
Note that this are the China prices with high markup due to export controls etc.
The price of a H800 80GiB in the US is today more like ~32k$USD .
But for using H800 clusters well you also need as fast as possible interconnects, enough motherboards, enough fast storage, cooling, building, interruption free power etc. So the cost of building a "H800" focused Datacenter is much much higher then multiplying GPU cost by number.
Still $400m seem unlikely.
/s TINFA -> this is not financial advice
OK, but does this quant fund have this amount of a spare resources to take a flyer on a vanity project?
Some variants of DeepSeek-R1 can be run on 2x H100 GPUs, and some people managed to get still quite decent results with a even stronger distilled mode running it on consumer hardware.
For DeepSeek-V3 even with 4bit quantization you need more like 16x H100.
"Assuming the rental price of the H800 GPU is $2 per GPU hour, our total training costs amount to only $5.576M. Note that the aforementioned costs include only the official training of DeepSeek-V3, excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data."
V3 was released a bit a month ago, V3 is not what took the world by storm but R1. The price everyone is talking about is the price for V3.
This is still quite impressive, given most people are likely to buy cloud infrastructure from AWS or Azure than build their own datacenter. So the Math checks out.
I don't think compute capacity built already will go waste, likely more and bigger things will get built in the coming years so most of it will be used for that purpose.
If they paid $70,000 per GPU[2] plus $5000 per 4-GPU compute node (random guess), then the hardware would have cost about $150M to build. If you add in network hardware and other data-centery-things, I could see it reaching into the $200M range. IMO $400M might be a bit of a stretch but not too wildly off base.
To reach parity with the rental price, they would have needed to re-train 70 times (i.e. over 12 years). They obviously did not do that, so I agree it's a bit unfair to cost this based on $2M in GPU rentals. Why did they buy instead of rent? Probably because it's not actually that cheap to get 2048 concurrent high-performance connected GPUs for 60 days. Or maybe just because they had cash for capex.
1: https://stratechery.com/2025/deepseek-faq/
2: https://www.tomshardware.com/news/price-of-nvidia-compute-gp...
Calling the training load for DeepSeek 6% of the value of that cluster seems generous. It probably used less of the recoverable value than that.
I think the salient point here is that the "price to train" a model is a flashy number that's difficult to evaluate out of context. American companies list the public cloud price to make it seem expensive; Deepseek has an incentive to make it sound cheap.
The real conclusion is that world-class models can now be trained even if you're banned from buying Nvidia cards (because they've already proliferated), and that open-source has won over the big tech dream of gatekeeping the technology.
The number does not include cost for personell, experiments, data preparation, chasing dead ends, and most importantly, it does not include the reinforcement learning step that made R1 good.
Furthermore, it is not factored in that both R3 and V1 are build on top of an enormous amount of synthetic data the was generated by other LLMs.
It increased the noise in the AI space by orders of magnitude. Every media outlet is bombarding you with a relentless torrent of half-true information, exaggered interpretation of single facts and speculation.
Sometimes I wonder whether this is amplified by a state actor?
Also curious, how Minimax-01, which is also an excellent model with impressive improvements, went by completely unnocited.
The only good thing is that this certainly put an end to OpenAIs price gauging - pretty sure that $200/month individual plan is not the limit of their imagination.
And that's somehow supposed to be end of NVIDIA? Hello?
https://openai.com/index/scaling-laws-for-neural-language-mo...
And a return to more normal world where progress comes from refinement of algorithms and approach rather than brute force compute.
Deepseek is just the poster child.
Thermodynamic neural networks may also basically turn everything on its ear, especially if we figure out how to scale them like NAND flash.
If anything, I would estimate that this is a space-race type effort to “win” the AI “wars”. In the short term, it might work. In the long term, it’s probably going to result in a massive glut in accelerated data center capacity.
The trend of technology is towards converging with or doing better than natural processes, not doing it 100000x less efficiently. I don’t think AI will be an exception.
If we look at what is -theoretically- possible using thermodynamic wells, with current model architectures, for instance, we could (theoretically) make a network that applies 1t parameters in something like 1cm2. It would use about 20watts, back of the napkin, and be able to generate a few thousand T/S.
Operational thermodynamic wells have already been demonstrated en silica. There are scaling challenges, cooling requirements, etc but AFAIK no theoretical roadblocks to scaling.
Obviously, the theoretical doesn’t translate to results, but it does correlate strongly with the trend.
So the real question is, what can we build that can only be done if there are hundreds of millions of NVIDIA GPUs sitting around idle in a few years? Or alternatively, if those systems are depreciated and available on secondary markets?
What does that look like?
Bigger models are more capable, but smaller models can be iterated on faster. For a couple months now most of the impressive achievements have been in increasingly smaller and cheaper models. Deepseek just has the perfect storm of impressive results, accessibility and international rivalry that made it go viral
It's actually the beginning of test time scaling. R1 has shown that a very simple reinforcement learning scheme can be used to teach the model how to think in a chain-of-though as an emergent property.
No addition pretraining data needed! Only more compute.
There was no market reaction when Deepseek v3 was released.
There was a massive market reaction when a Deepseek app was released.
There is too much damn irrationality in the AI and AI adjacent market right now, and pop articles like this are only making it worse.
Jevons Paradox absolutely holds for model development, and Deepseek should (and is - all my cybersecurity startup peers are investigating the feasibility of building their own models now) be viewed as an opportunity for smaller or medium stage companies to build their own competitive domain specific models, while not having to pay what is essentially protection money to OpenAI or Anthropic.
The only people at risk with the democratization of model development are the investors in foundational model companies
---------
Also, I hope Deepseek FINALLY reprioritizes distributed systems education in CS.
There are too many "ML Engineers" who do not know the basics of OS or Hardware Engineering (eg. Cannot optimize Mellanox hardware, really dig into workload optimization, etc) and are thus burning resources inefficiently. Most MLEs I meet now just import a package and glue code together WITHOUT also understanding the performance capabilities that their compute has.
Most universities in the US either don't require OS or CompArch classes for CS majors (eg. Harvard) or dumb them down significantly (eg. Not going deep into Linux kernel implementation, scheduling, etc).
This is why the entire cybersecurity industry has largely shifted to Israel and India, and the skillset overlaps significantly with MLOps and MLEng (if you know how to implement spinlocks in C you can easily be taught CUDA)
<"Old Man yelling at clouds" rant over>
It is very possible, when things require advanced technical knowledge and experience
><"Old Man yelling at clouds" rant over>
They are too busy deploying ruby apps and juggling jsons over https to mess with ugly cpp, segfaults, kernel details, hardware intrinsics and semiconductors
> They are too busy deploying ruby apps and juggling jsons over https
Ruby is "legacy" as well nowadays. Most younger devs I meet tend to really only understand JS and Python, and that too while heavily relying on outside packages or dependencies.
Not a bad thing per say, but if a VC like me has deeper knowledge about OS internals or Mellanox tuning than some of the (American) MLEs in companies they've done due diligence on, something's very wrong with the talent pipeline.
It is why empires collapse from miscalculations about war by their experts, it is the foundational mental model underlying communism that simply cannot work but keeps being attempted, and it is how countries can be destroyed in every which way while the experts maintain that everything is just fine.
Never underestimate the powers of the ego, the destroyer of worlds. I mean they don’t know what they are doing, but we have advanced technical knowledge and experience, so we should clearly be in charge of all things, including those beyond our narrow scope of advanced technical knowledge and experience.
No communism doesn't work either. I wish there was more alternatives.
And some of the usual examples of "socialism doesn't work" might have worked if US hadn't interfered. E.g. in Chile, no system would have survived US actively supporting Pinochet. We will never know how socialism would have played out there if US (and Soviet too) left it alone.
UCB's CS162 - https://cs162.org/
MIT 6.1810 - https://kaashoek.github.io/65810-2023/
These were tablestake courses that teach you the basics of OS and Systems Development (eg. Synchronization, Scheduling, etc).
After you understand that, then you'd start digging into HPC classes AND then ideally go into ML.
At least, this was the workflow and mental model I followed a decade+ ago.
And this is quite separate from the business of it. For example Jevons paradox is quite apparent in the airline industry, but airlines are notorious for possibly never having net made a profit because of high capex and aggressive pricing competition. Jevons paradox is not any reason at all for a runaway valuation of a company.
Furthermore, it's Nvidia that owns Mellanox - which is the owner of the Infiniband IP which is used in just about every DC or cluster, because no other vendor came as close for interconnect performance
That said, the reason I don't invest Nvidia is that they're selling one more thing - AI legitimacy. Every tech shop is buying Nvidia because all their CEOs need to go to the board and say, we've got so and so AI strategy. Sure as hell Meta isn't spending 50 billion a year because AI makes them so much damn money (supposedly ML saved their ass when Apple put on App Tracking Transparency, but 99% suds it's not the ML that needs a gigawatt GPU cluster to run). And the very real calculus for investors/CEOs is that if DeepSeek made this on a thousandth of the budget (conflicting reports but high probably DS has a fraction of the compute), whYs stopping Zuckerberg from taking 10B of that capex spend and offering every name on that DeepSeek paper a 10 million salary?
Nobody knew what it was yesterday and now everyone is using it as if they came up with the idea themselves
https://trends.google.com/trends/explore?date=today%205-y&q=...
I see two pages, just of submissions on HN, before early last year:
https://hn.algolia.com/?dateEnd=1706400000&dateRange=custom&...
I see over 900 results just for comments:
https://hn.algolia.com/?dateEnd=1706400000&dateRange=custom&...
I personally have brought it up when it's relevant, like when people think energy efficiency improvements always make total use go down:
https://news.ycombinator.com/item?id=14764845
https://news.ycombinator.com/item?id=19543772
And when a more efficient website showed more load on the server (because it widened usage to a broader audience):
https://news.ycombinator.com/item?id=13602792
How about the more charitable hypothesis: It's a relevant dynamic because this is a situation where the paradox most strongly applies: efficiency improvement in use of an input whose output has insatiable demand. And so people bring it up, even if they just recently learned it from another comment.
It's okay to say something true ... even if someone else already said it elsewhere.
It's the fancy big word of the day if you want to appear smart on social media.
Nvidia currently has ~75% margins and ~90% growth. If either of those nobs get turned down slightly, their valuation can tank.
Which is what happened.
If Nvidia's 10-year growth forecast went from ~40% per year (29x in 10 years) to ~30% per year (14x in 10 years) - that's a 50% smaller future company you're expecting.
It's still great for Nvidia. It's just LESS great than people thought before, which means their market cap comes down a ton.
That's what you're all forgetting.
What margins do you think Nvidia is charging for RTX which getting more and more expensive every generation?
What margins do you think Project Digits will be for Nvidia? That thing costs 50% more than a RTX card and on the RTX card Nvidia only earns on the chip which is only part of the card. Digits is a 100% Nvidia product.
Does it matter if 1 Blackwell DC GPU or 13 Digits are sold? No, Nvidia has probably an insane margin on Digits and has only one goal, total spread and integration of Nvidia HW in any AI workload. Because then Nvidia can offer SW with super margins.
Think of Nvidia Enterprise AI, Omniverse, Clara, Isaac, DriveSim and many more. If 95% of AI acclerators are from Nvidia then it's only a small step to also use Nvidia SW frameworks.
Seeing that DeepSeek-R1 has benchmark scores like o1 is not the same as seeing that people actually like it, that it's being adopted, and that stated training costs are getting accepted as credible.
(But I agree that there is a lot of irrationality in these markets.)
> There was no market reaction when Deepseek v3 was released.
> There was a massive market reaction when a Deepseek app was released.
This signals that there was money on the table by paying attention to the underlying tech.
Do you know how to optimize the MTU of an Infiniband interconnect? Do you understand how to schedule multiple models being trained concurrently? Do you know why you cannot directly leverage bare metal compute within a Docker image?
This is important Infra knowledge you need to take full advantage of any model you are training on your own hardware. And this is why Deepseek was successful - they understood the ins-and-outs of systems programming and the H800 architecture to maximize the compute performance they needed to train their model.
Nothing!
That's my point! I've met a number of "MLEs" who couldn't push back like that with my very basic "fizzbuzz" question
There is a wasteland of blogspam around the topic of getting started with ML and it’s hard to know what is useful at the beginning.
Then just start tinkering. I got interested in ML because of sport's analytics and betting markets so I read a lot of papers on that topic and books similar to Bayesian Sports Models in R by Andrew Mack[1].Also, Jake VanderPlas's Python Data Science Handbook is good[2].
Ideally, find a vertical you're interested in where experts have applied ML and read their papers/books and work backwards from there.
[0]: https://statquest.org/statquest-store/ [1]: https://www.goodreads.com/book/show/216487475-bayesian-sport... [2]: https://www.oreilly.com/library/view/python-data-science/978...
China is good on dumb stupid manufacturing. It does this with so much enthusiasm HNers are left scratching their heads why the excitement. The possibility of China manufacturing the next computing powerhouse was too high. Higher than it could ever be in the US.
The threat to US leadership was that the fastest supercomputer for AI in the world could be in China (the first sputnik moment per Barack Obama). So the chip embargo was to prevent that from happening.
So China did not spend/waste $$ & time building hardware. They spent it on software. Normally the Chinese version is cheap enough and good enough. But this time it was cheap enough and surprisingly terrific enough. Props to Liang the founder for having that game face on.
It seems more like , investors were looking for a significant event to pop the obvious bubble if the last year
Investors who were actually invested, or onlookers with FOMO who hadn't invested yet?
While I think the "AI" companies invested in are overpriced*, I think the market as a whole hasn't even begun to price in the changes in store across non-tech sectors and industries even if LLMs were to advance no further than now, with all additional progress at the prompting and agentic levels.
* Even here, Jevons paradox applies. So what's being called "priced for perfection" may actually be priced for business as usual with multiplied demand.
What? Oh, come on now.
ftfy