I've tried their 7b model, running locally on a 6gb laptop GPU. Its not fast, but the results I've had have rivaled GPT4. Its impressive.
People who can use the 585B model will use the best model they can have. What DeepSeek really did was start an AI "space race" to AGI with China, and this race is running on Nvidia GPUs.
Some hobbyists will run the smaller model, but if you could, why not use the bigger & better one?
Model distillation has been a thing for over a decade, and LLM distillation has been widespread since 2023 [1].
There is nothing new in being able to leverage a bigger model to enrich smaller models. This is what people that don't understand the AI space got out of it, but it's clearly wrong.
OpenAI has smaller models too with o1 mini and o4 mini, and phi-1 has shown that distillation could make a model 10x smaller perform as well as a much bigger model. The issue with these models is that they can't generalize as well. Bigger models will always win at first, then you can specialize them.
Deepseek also showed that Nvidia GPUs could be more memory-efficient, which catapults Nvidia even further ahead of upcoming processors like Groq or AMD.
Although not all commodities will work like fossil fuels did in Jevon’s Paradox. It could be the case that demand for AI doesn’t grow fast enough to keep demand for chips as high as it was, as efficiency improves.
We tried that, though. NPUs are in all sorts of hardware, and it is entirely wasted silicon for most users, most of the time. They don't do LLM inference, they don't generate images, and they don't train models. Too weak to work, too specialized to be useful.
Nvidia "wins" by comparison because they don't specialize their hardware. The GPU is the NPU, and it's power scales with the size of GPU you own. The capability of a 0.75w NPU is rendered useless by the scale, capability and efficiency of a cluster of 600w dGPU clusters.
GPUs will continue to be bought up as fast as fabs can spit them out.
Couldn’t you say that about Blackwell as well? Blackwell is 25x more energy-efficient for generative AI tasks and offer up to 2.5x faster AI training performance overall.
What does that tell us?
The industry is compute starved and that makes totally sense.
The tranformer model on which current LLMs are based on are 8 years old. But why took it so much time to get to the LLMs only 2 years ago?
Simple, Nvidia first had to push the compute at scale strongly. Try training GPT4 on Voltas from 2017. Good luck with that!
Current LLMs are possible thanks to the compute Nvidia has provided in the past decade. You could technically use 20 year old CPUs for LLMs but you might need to connect a billion of them.
You can rent 10k H100 for 20 days with that money. Go and knock yourself out because that compute is probably higher than what DeepSeek received for that money. And that is public cloud pricing for single H100. I'm sure if you ask for 10k H100 you'll get them at half price so easily 40 days of training.
DeepSeek has fooled everyone by telling them that they need only so less money and people think that they only need to "buy" $5M worth of GPU but that's wrong. The money is the training costs of renting the GPU training hours.
Somebody had to install the 10k GPUs and that's paying $300M to Nvidia.
if there's evidence to the contrary I'd love to see. in any case I don't think a h800 is even 20x better than a h100 anyway, so the 20x increase has to be wrong.
Also, everything we know about LLMs points to an entirely predictable correlation between training compute and performance.
High difficulty:
id = 37810
word = dendroid
pos = noun
sense = (mathematics) A connected continuum that is arcwise connected and hereditarily unicoherent.
elo = 2408.61936886416
sentence2 = The dendroid, that arboreal structure of the Real, emerges not as a mere geometric curiosity but as the very topology of desire, its branches both infinite and indivisible, a map of the unconscious where every detour is already inscribed in the unicoherence of the subject's jouissance.
Low difficulty: id = 11910
word = bed
pos = noun
sense = A flat, soft piece of furniture designed for resting or sleeping.
elo = 447.32459484266
sentence2 = The city outside my window never closed its eyes, but I did, sinking into the cold embrace of a bed that smelled faintly of whiskey and regret.It's supposed to. There was an info that the longer length of 'thinking' makes o3 model better than o1. I.e. at least at inference compute power still matters.
compute matters, but performance doesn't scale with compute from what I've heard about o3 vs o1.
you shouldn't take my word for it - go on the leaderboards and look at the top models from now, and then the top models from 2023 and look at the compute involved for both. there's obviously a huge increase, but it isn't proportional
Similarly, as fast as processors have gotten, people still complain their applications are slow. Because they do so much more.
Generally applicable ML is still in its infancy, and usage is exploding. All those newfound spare cycles will get soaked up fairly quickly.
Blackwell DC is $40k per piece and Digits is $3k per piece. So if 13x Digits are sold then it's the same turnover as a DC GPU for Nvidia. Yes, maybe lower margin but Nvidia can easily scale digits into masses compareds to Blackwell DC GPUs.
In the end, the winner is Nvidia because Nvidia doesn't care if DC GPU, Gaming GPU, Digits GPU, Jetson GPU is used for AI as long as Nvidia is used 98% of time for AI workloads. That is the world domination goal, simple as that.
And that's what Wallstreet doesn't get. Digits is 50% more turnover than the largest RTX GPU. On average gaming GPU turnover is probably around $500 per GPU. Nvidi probably sells 5 million gaming GPUs per quarter. Imagine they could reach such amounts of Digits. That would be $15b revenue and almost half of current DC revenue with Digits only.
Electricity demands will plummet when transistors take the place of vacuum tubes.
Anything other than their 671b model are just distilled models on top of Qwen and Llama using their 671b reasoning data output, right?
If only I could figure out how to buy NV stock quickly before it rebounds
Distilled models are nothing new.
Frontier models are heavily compute constrained - the leading AI model makers have got way more training data already than they could do anything with. Any improvement in training compute-efficiency is great news for them, no matter where it comes from. Especially since the DeepSeek folks have gone into great detail wrt. documenting their approach.
Citation needed.
Also current SOTA models are good enough that you can generate endless training data by letting the model operate stuff like a C compiler, python interpreter, Sage computer algebra, etc.
Could be entirely wrong here - would love a fact-check by industry insider or journalist.
Making training more effective makes every unit of compute spent on training more valuable. This should increase demand unless we've reached a point where better models are not valuable.
The openness of DeepSeek's approach also means that there will be more smaller entities engaging in training rather than a few massive entities that have more ability to set the price they pay.
Plus reasoning models substantially increase inference costs, since for each token of output you may have hundreds of tokens of reasoning.
Arguments on the point can go both ways, but I think on the balance I would expect any improvements in efficiency increase demand.
2. Even if DeepSeek's budget claims are true, they trained their model on the outputs of an expensive foundation model built from a massive capital outlay. To truly replicate these results from scratch, it might require an expensive model upstream.
Given they've reproduced earlier model's and vetted it - I think it's probably safe to assume that these new models are not out of thin air - but until somebody reproduces it, it's up in the air.
Nvidia is growing profits faster than income.
Nvidia's net profit margin is 55% (vs Apple 15%) and they have an operating income of $21B vs Apple's $29.5
These are some pretty impressive financial results - those growth rates are the reason people are bullish on it.
How sound is the investment thesis when a bunch of online discussions about a technical paper on a new model can cause a 20% overnight selloff? Does Apple drop 20% when Samsung announces a new phone?
If it were valued that way, the P/E would be over 100.
Feel free to say Nvidia is overvalued, but you have to get the financials right.
That probably won't be the first question we ask AGI if/when we ever get there, but it will be near the top of the list.
What needed 1000k of Voltas, needed 100k of Amperes, needed 10k of Hopper, will need 1k of Blakwell.
Nvidia has increased compute by a factor of 1 million in the past decade and it's no where near enough.
Blackwell will increase training efficiency in large clusters a lot compared to Hopper and yet it's already sold out because even that won't be enough.
No one expects this growth to be sustained for a decade. Companies aren't prices based on hypothetical growth rates in 10 years time.
As you see NVidia doesn't stand out much, it's even lower than Amazon.
And NVDA’s P/E benefits from very recent huge spending that may not continue.
Look at their PEG ratios.
In theory it’s more about forward profits per share, taking into account growth over many years. And Nvidia is growing faster than any company with that much revenue.
Obviously the future is hard to predict, which leaves a lot of wiggle room.
But I say in theory, because in practice it’s more about global liquidity. It has a lot to do with passive investing being so dominant and money flows.
Money printer goes brrr and stonks go up.
That is not the only thing that matters, but it seems to be the main thing.
If it were really about future profits most of these companies would long since be uninvestable. The valuations are too high to expect a positive ROI.
DeepSeek supposedly nullifies that last part.
I can't see how DeepSeek hurts Nvidia, if Nvidia is what enables DeepSeek.
the simplest way to present the counter argument is:
- suppose you could train the best model with a single H100 for an hour. would that hurt or harm nvidia?
- suppose you could serve 1000x users with a 1/1000 the amount of gpus. would that hurt or harm nvidia?
the question is how big you think the market size is, and how fast you get to saturation. once things are saturated efficiency just results in less demand.
I'm skipping over some details of course, but the current Nvidia valuation, or rather the valuation a few days ago, was based on them being the only company capable of producing chips that can train the best models. That wasn't true for those in the know before, but is now very much more clearly not true.
They are getting a lot of money, but their stock price is in a completely different universe. Not even that $500G deal people announced, if spent exclusively on their products could justify their current price. (Nah, notice that just the change on their valuation is already larger than that deal.)
https://www.reddit.com/r/LocalLLaMA/comments/1c0je6h/comment...
"The biggest threat to NVIDIA is not AMD, Intel or Google's TPU. It's software. Sofware eats the world!"
"That's what software is going to do. A new architecture/algorithm that allows us current performance with 50% of the hardware, would change everything. What would that mean? If Nvidia had it in the books to sell N hardware, all of a sudden the demand won't exist since N compute can be realized with the new software and existing hardware. Hardware that might not have been attractive like AMD, Intel or even older hardware would become attractive. They would have to cut their price so much, the violent exodus from their stocks will be shocking. Lots of people are going to get rich via Nvidia, lots are going to get poor after the fact. It's not going to be because of hardware, but software."
A lot of people are saying that I'm wrong on other hardware like AMD or Intel, but this article by Stratechery agrees, all other hardware vendors are possibly relevant again. I didn't talk about Apple because I was focused on the server side, Apple has already won the consumer side and is so far ahead and waiting for the tech to catch up to it.
The biggest threat to Nvidia is still more software optimization.
Today, it's simple. Apple has 25% unit share in smartphone markets and 75% profit share. Apple makes 3x the profit of ALL OTHER smartphone vendors combined.
And this is exactly where Nvidia's goal is. The AI compute market will grow, Nvidia will lose unit market share but Nvidia will retain their profit market share. Simple as that.
And by the way, Nvidia is way ahead in SW compared to alternatives. Most here have the DIY glasses on. But enterprises and businesses have different lenses. For those not being Tech they need secure and working solution with enterprise grades. Nvidia is among the few to offer this with Enterprise AI solutions (NeMo, NIMs, etc.). Nvidia's SW moat isn't CUDA, CUDA is an API for performance and stability. Nvidia's SW moat is in the frameworks for applications for many differnt industries and of course ALL Nvidia SW will require Nvidia HW.
A company using Nvidia enterprise SW solutions and consultancy will never use anything except Nvidia HW. Nvidia has a program with >10k AI startups being supported with free consulting and HW support. Nvidia is basically grooming their next generation customers by themselves.
You have no idea, many think Nvidia is only selling some chips and that's where they are wrong. Nvidia is a brand, an ecosystem and they will continue to grow from there. See gaming, much more standards and commodity in SW than AI SW. There is no CUDA, you can swap a Nvidia card with AMD card within a minute. So let me know, how come for 2 decades that Nvidia has continously 80-95% market share?
NVDA Net income, Quarter ending in ~Oct2024: $19B. AMD? $771M. INTC? -$16.6B. QCOM? $3B. AAPL? $14B.
Revenue growth, YoY? +93%. AMD? +17%. INTC? -6%. QCOM? +18%. AAPL? +6%.
Margin? 55%. AMD? 11%. INTC? -125%. QCOM? 28%. AAPL? 15%.
P/E Ratio? 46. AMD? 103. INTC? N/A, unprofitable. QCOM? 19. AAPL? 34. NFLX? 54. GME? 151.
Their P/E Ratio doesn't even classify them as all that overvalued. Think about that. Price to earnings, they are cheaper than Netflix, Gamestop, they're about the same level as WALMART, you know, that Retailer everyone hates that has practically no AI play, yeah their P/E is 40.
Nvidia is an insane company. Insane. We've had three of the largest country-economies on the planet announce public/private funding to the tune of 12 figures, maybe totaling 13 figures when its all said and done, and NVDA is the ONLY company on the PLANET that sells what they want to buy. There is no second player. Oh yeah, Google will rent you some TPUs, haha yeah sure bud. China wants to build AI data centers, and their top tech firms are going to the black market smuggling GPUs across the ocean like bricks of cocaine rather than rely on domestic manufacturers, because not even other AMERICAN manufacturers can catch up.
Sure, a 10x drop in cost of intelligence is initially perceived as a hit to the company. But, here's the funny thing about, let's say, CPUs: The Intel Northwood Pentium 4 was released in 2001; with its 130nm process architecture, it sipped a cool 61 watts of power. With today's 3nm process architecture, we've built (drumroll please) the Intel Core Ultra 5 255, which consumes 65 watts of power. Sad trombone? Of course not; its a billion times more performant. We could have directed improvements in process architecture toward reducing power draw (and certainly, we did, for some kinds of chips). But, the VAST, VAST, VAST majority of allocation of these process improvements was in performance.
The story here is not "intelligence is 10x cheaper, so we'll need 10x fewer GPUs". The story is: "Intelligence is 10x cheaper, people are going to want 10x more intelligence."
A credible lab making a credible claim to massive efficiency improvements is a credible threat to Nvidia's future earnings. Hence the stock got sold. It's not more complicated than that.
Its obviously constrained by this hardware and this model size as it does some strange things sometimes and it is slow (30 secs to respond) but I've got it to do some impressive things that GPT4 struggles with or fails on.
Also of note I asked it about Taiwan and it parroted the official CCP line about Taiwan being part of China, without even the usual delay while it generated the result.
Inference cost - DeepSeek is charging less than OpenAI to use its public API, but that isn't an indicator of anything since it doesn't reflect the actual cost of operation. It's pretty much a guarantee that both companies are losing money. Looking at DeepSeek's published models the inference cost is in the same ballpark as Llama and the rest.
Which leaves training, and that's what all the speculation is about. The CEO said that the model cost $5.5M and that's what the entire world is clinging on. We have literally no other info and no way to verify it (for now, until efforts to replicate it start to show results).
Again, the weights are public. You can run the full-fat version of R1 on your own hardware, or a cloud provider of your choice. The inference costs match what DeepSeek are claiming, for reasons that are entirely obvious based on the architecture. Either the incumbents are secretly making enormous margins on inference, or they're vastly less efficient; in the first case they're in trouble, in the second case they're in real trouble.
Deepseek has distilled deepseek R1 into a couple of smaller open source models, but neither R1 or v3 are distilled themselves.
If meme stocks were imploding, why is Tesla fine?
This is about DeepSeek.