Nvidia’s $589B DeepSeek rout
finance.yahoo.com
finance.yahoo.com
If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.
I'd say the wider industry could learn a thing or two, but as other commentors have joked. The line must go up
Creativity from constrained minds?
Reminds me of Japanese cars and the OPEC boycott in the 1970s...
https://en.wikipedia.org/wiki/Standard_Portable_Intermediate...
Or even better in the unified language like SYCL?
https://cdrdv2-public.intel.com/786536/Heidelberg_IWOCL__SYC...
You can do this just fine in CUDA, no PTX required. Of course all the major shops are using inline PTX at the very least to access the Tensor cores effectively.
Whenever I prompt: "Do not do anything"
It always does <something>.
Yep. A lot of times, the responses I get remind me of Simone in Ferris Bueller's Day Off: https://www.youtube.com/watch?v=swBtLPWeKbU
If you end up making a new model, please teach it that less is more and call it "LAIconic".
If line keeps going up, line does catastrophic or potentially apocalyptic harm, given our current circumstances.
they claim it's able to run models with 200B parameters on a single node and 400B when paired with another node
Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year? It was always the second one, and doing the training cheaper makes it even more so.
But this implies that we could use those same resources to train even bigger models, right? Except that you then have the same problem. You have a bigger model, maybe it's better, but if you've made inference cost linearly more because of the size and the size is now 40x bigger, you now need that much more compute for inference.
I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4.
I've just switched to it for my local inference.
From DeepSeek-R1 paper:
> As shown in Table 5, simply distilling DeepSeek-R1’s outputs enables the efficient DeepSeekR1-7B (i.e., DeepSeek-R1-Distill-Qwen-7B, abbreviated similarly below) to outperform nonreasoning models like GPT-4o-0513 across the board.
and
> DeepSeek-R1-14B surpasses QwQ-32BPreview on all evaluation metrics, while DeepSeek-R1-32B and DeepSeek-R1-70B significantly exceed o1-mini on most benchmarks.
and
> These [Distilled Model Evaluation] results demonstrate the strong potential of distillation. Additionally, we found that applying RL to these distilled models yields significant further gains. We believe this warrants further exploration and therefore present only the results of the simple SFT-distilled models here.
I run my models in an agentic framework with fast models that can ask slower models or APIs when needed. It works perfectly, 60 percent of the time lol.
https://mlnotes.substack.com/p/the-valleys-going-crazy-how-d...
The only reason why a team would allocate time on memory optimizations and writing NVPTX code rather than focusing on posttraining is if they severely struggled with memory during training.
I mean, take a look at the numbers:
https://www.fibermall.com/blog/nvidia-ai-chip.htm#A100_vs_A8...
This is a massive trick pulled by Jensen, take the H100 design whose sales are regulated by the government, make it look 40x weaker and call it H800, while conveniently leaving 8-bit computation as fast as H100. Then bring it to China and let companies stockpile without disclosing production or sales numbers, and have no export controls.
Eventually, after 7 months, US govt starts noticing the H800 sales and introduces new export controls, but it's too late. By this point, DeepSeek has started research using fp8. They slowly build bigger and bigger models, work on the bandwidth and memory consumptions, until they make r1 - their reasoning model.
Tech or politics related, he's off the deep end.
He's a lucky mensch, no more, no less.
At some point, the models _have_ to do "continuous integration" to provide the "AGI" that's wanted out of this tech.
The only thing it's not good for is the idea that OpenAI and/or Anthropic will eventually become profitable companies with market caps that exceed Apple's by orders of magnitude. Oh no, anyway.
Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to happen in the long run?
https://en.m.wikipedia.org/wiki/Knowledge_distillation
So the takeaway is they have no moat
By releasing a lot of it opensource everyone has their hands on it. Opens the door to new companies.
Or a simple mental model, there has been this ability for third parties to get quite close to leading frontier models. The leading frontier models takes hundreds of millions of dollars and if someone is able to copy it within a years time for significantly less capital, its going to be hard game of cat and mouse.
That said, you have to distinguish between "good for the field of AI, the AI industry overall, and users of AI" from "good for a couple of companies that want to be the sole provider of SOTA models and extract maximum value from everyone else to drive their own equity valuations to the moon". Deepseek is positive for the former and negative for the latter.
I don't really understand why the stock market has decided this affects nvidia's stock price though.
The jury is still out on how much improvement DeepSeek made in terms of training and inference compute efficiency, but personally I think 10x is probably the actual improvement that's being made
But in business/engineering/manufacturing/etc if you have 10x more efficiency, you're basically going to obliterate the competitions.
>output quality has ceased to be compute-bound
You raised an interesting conjecture and it seems that it's very likely the case.
I know that it's not even a full two years that ChatGPT-4 has been released but it seems that it take OpenAI a very long time to release ChatGPT-5. Is it because they're taking their own sweet time to release the software not unlike GIMP, or they genuinely cannot justify the improvement to jump from 4 to 5? This stagnation however, has allowed others to catch up. Now based on DeekSeek claims, anyone can has their own ChatGPT-4 under their desk with Nvidia project Digits mini PCs [1]. For running DeepSeek, 4 units mini PCs will be more than enough of 4 PFLOPS and cost only USD12K. Let's say on average one subscriber user pays OpenAI monthly payment of USD$10, for 1000 persons organization it will be USD$10K, and the investment will pays for itself within a month, and no data ever leave the organization since it's a private cloud!
For training similar system to ChatGPT-4 based on DeepSeeks claims, a few millions USD$ is more than enough. Apparently, OpenAI, Softbank and Oracle just announced USD$500 Billions joint ventures to bring the AI forward with the new announced Stargate AI project but that's 10,000x money [2],[3]. But the elephant in the room question is that, can they even get 10x quality improvement of the existing ChatGPT-4? I really seriously doubt it.
[1] NVIDIA Puts Grace Blackwell on Every Desk and at Every AI Developer’s Fingertips:
https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwe...
[2] Trump unveils $500bn Stargate AI project between OpenAI, Oracle and SoftBank:
https://www.theguardian.com/us-news/2025/jan/21/trump-ai-joi...
[3] Announcing The Stargate Project:
The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory.
Fortune 100 companies will still want the biggest toolshed to invent the next paradigm or to be the first to get to AGI.
What is "deepseek specification"? Deepseek was trained on NVDA chips. If chinese vendors could build chips as good as NVDA it wouldn't have such a dominant position already, that hasn't changed
The claimed training breakthrough is an optimization targeting NVidia chip, not something that reduces NVidia's relative advantage. Even if it is easily generalizable to other vendors hardware, it doesn't reduce NVidia's advantage over other vendors, it just proportionately scales down the training requirements for a model of a given capacity. Which, maybe, very short term reduces demands from the big existing incumbents, but it also increases the number of players for which investing in GPUs for model training at all is worthwhile, increasing aggregate demand.
But your point is well taken and perhaps both mine and GP's metaphors break down.
Either way, we saw massive spikes in demand for Nvidia when crypto mining became huge followed by a massive drop when we hit the crypto winter. We saw another massive spike when LLMs blew up and this may just be the analogous drop in demand for LLMs
Cue all of China rushing to Jensen to buy all the H800s they can before the embargo gets tightened, now that their peers have demonstrated that they're useful for something.
At least briefly, Jensen's customer audience increased.
I still think this is bullish: more people will be buying chips once cheaper and more accessible, and the things the will be training with be 1,000% to 10,000% larger
(Many individual people are already saying that, but they aren't the people buying the GPUs for this in the first place. Steam engines weren't universally popular either when they were introduced to society.)
Its just that it costs too much to do that for the hoi polloi who think everything digital should be free forever.
Really? Has anyone made a useful, commercially successful product with it yet?
People talk like the end user demand part of the equation is really solved when invoking an Econ 101 magical interpretation of the Law of Supply and Demand or Jevons Padadox.
Deep integration into iOS won't be tacked on in a rush to market addition to OS.
But to be clear, none of these things are commercial products. They're gimmicks. Google is an ads company, they make their money selling ads. Apple is a computer company, they make their money selling computers. These "AI products" are a circus sideshow.
Aren't millions or even tens of millions students using ChatGPT for example? To me that sounds like a commercial success (and looks comparable with the usage of Google Search - a money printing machine for more almost 30 years now - in the first years)
And enterprise-wise - heard recently a VP complaining about entering expenses. As we don't have secretaries anymore in the civilized world, that means "Agents AI" is going to have a blast.
(i'm long on NVDA and wondering is it enough blood on the streets to buy more :)
It isn't because it's not making them any money. Having users doesn't mean you have a business. If you sell two dollars for one dollar having more users is not a blessing financially. Of course you could slap ads on it, like Google, but unlike Google openAI has no moat and there's already ten competitors. Competition eliminates profit and AI is being commoditized faster than pretty much anything else.
what was Google moat?
>and there's already ten potential competitors. Competition eliminates profit and AI is being commoditized faster than pretty much anything else.
we're discussing NVDA. Where are its competitors? ChatGPT having 10 competitors only makes things better for NVDA.
>Competition eliminates profit
Competition weeds out bad/ineffective performers which is great. History of our industry is littered with competition taking out bad performers, and our industry is only better for that. Fast commoditization of AI is just great and fits the best patterns like say PC-revolution (and like it i think the AI-revolution wouldn't be just one app/user-case, it will be a tectonic shift instead).
I read somewhere that OpenAI brought in $3.7 billion in 2024, and made a loss of $6 billion. So... no I don't think that is an example. They want to make a commercially successful product, but ChatGPT doesn't seem to be there yet.
> And enterprise-wise - heard recently a VP complaining about entering expenses. As we don't have secretaries anymore in the civilized world, that means "Agents AI" is going to have a blast.
We don't have secretaries because the word became unfashionable. They are called PAs or executive assistants or something like that now. They're still there, but if anything the need for them has probably been reduced with (non-AI) computers (calendars, contacts, emails, electronic documents, etc.) so I'm not sure that there is some enormous unmet demand for them.
The question is whether ChatGPT (the product) and running thr API is profitable or at at least whether the trend is that cost are coming down.
The same is happening in enterprise tier products, Copilot 365 is still an extra SKU to count while Google Gemini Advanced has been integrated into the Workspace offering (i.e. they actually force you for an upsell of ~20% per user license for something we didn't ask, but I digress). At least that's a better alternative that paying +20 USD per license.
Prices need to and will go down, and business models will have to change and they are already doing so. But I'm not sure if OpenAI is really ready for that.
I gave it an easy one, “How many of the actors from the original Star Trek are still alive”. It gave me accurate information as of its training cut off date. But ChatGPT automatically did a web search to validate its answer. I had to choose the search option for it to look up later info.
With ChatGPT even when it doesn’t do a web search automatically, I can either tell it to “validate this” or put in my prompt “validate all answers by doing a web lookup”.
Then I gave it a simple compounding interest problem with monthly payments and wanted a month by month breakdown. DeepSeek used its multi step reasoning like o1 and was slower. ChatGPT 4o just created a Python script and used its internal Python interpreter.
Then DeepSeek started timing out.
This is the presentation of “what are some of the best places to eat in Chicago?”
https://chatgpt.com/share/6799510f-f4a8-8010-b80d-100c95d36d...
It doesn’t show on the shared link. But in the app it gives you a map of the restaurants or you can choose a list.
I can’t share a conversation with DeepSeek (?). But suffice it to say, the interface wasn’t as good.
I’m not saying the underlying technology of DeepSeek isn’t “good enough”. But the end user product is severely lacking.
Either way that means a lot more NVDA hardware being sold. You still need CUDAs as rocm is still not there yet. In fact NVDA needs to churn out more CUDAs than ever.
Not sure what is supposed to happen to the inference demand but I guess that could be modeled as more of a long-run thing, as inference is going to be very coin-operated (companies need it to be net profitable now) whereas training is more of a build now profit later game.
A good fundamental analysis is probably very hard to get right, and the game is probably just guessing which way everyone else will guess the herd is going to jump.
If every well funded start-up can have a shot, then they buy more GPUs and the big players will need to buy even more to stay noticeably ahead.
I argue it doesn't apply to generative AI because its outputs are mostly no good, have no utility, or are good but only in limited commercial contexts.
In the first case, a machine that produces garbage faster and cheaper doesn't mean demand for the garbage will increase. And in the second case, there aren't enough buyers for high-quality computer-generated pictures of toilets to meaningfully boost demand for Nvidia's products.
Yes, the value of those only exists mostly if your internal team is too stubborn to change its opinion. But that seems to be the norm. And the value (those) consultants add is not that high in the first place! They don't have the internal knowledge of _why_ things are fucked up _your particular way_ anyways. That part your team has to contribute anyhow. So the value add shrinks to "throw ideas over the wall and see what sticks". And LLMs are excellent at that.
Yes, that doesn't replace a highly technical consultant that does the actual implementation. Yes, that doesn't give you a good solution. But it probably gives you 5 starting points for a solution before you even finish googling which consultancy to pick (and then waiting for approval and hoping for a goodish team). And that's a story that I can map to reality (not that I like this new bit of information about reality..)
If we accept that story about LLM value, then I think NVIDIA is fine. That generated value is far greater than any amount of energy you can burn on inferring prompts and the only effect will be that the compute-for-training to compute-for-inference ratio decreases further
That exec was hiring consultant and no longer is, in meaningful proportion, thanks to LLM?
I don't know how things are in other white-collar industries (except wrt. creative jobs like copywriting and graphics design, where generative AI is even better at the job as it is at coding), but the incentives are similar so I expect most of the actual work is done by juniors anyway, and subject to replacement by models less sophisticated than people would like to imagine they need to be.
There are basically two reasons for consultants:
Either you need some once-removed thing (be that you selling/buying some part to/from a competitor, accounting, lawyering). That part sensible people will not replace by LLMs. But here it isn't that you yourself lack the expertise at all. Here it is absolutely necessary, that someone else does the actual implementation.
Or you have some general "we need to do better" feeling. And here you have again two options: 1) you know _what_ your problem is and you just need the best solution there is. This is (obviously somewhat tongue-in-cheek) essentially corporate espionage. Again you cannot replace that with an LLM (or maybe you can I don't know), but you will pay a lot of money for it. Or 2) you don't know what the problem is. Now you are competing in finding an appropriate starting point with 20-somethings fresh from university that are not wanted in your organisation and therefore won't get access to the relevant information anyhow. So yeah, I'm willing to believe that the typical success rate of those consultancy projects is low to negative.
If you are only given a 3 week crash course in $BUSINESS, you won't be able to produce much more than a generic set of "have you thought about that?". And THAT is something that I believe LLMs to be reasonably good at. And they are dirt cheap and instantly available compared to any kind of human consultant.
Now I don't think that will necessarily a net-negative for consultants in general. I do think, similar to ATMs, that consultations are mostly becoming cheaper by that.
But at some point, if the marginal product gets high enough, the world needs not as many, because money spent on other inputs/factors pays off more.
This is a classic problem with extrapolation. Making people more efficient through the use of AI will tend to increase employment... until it doesn't and employment goes off a cliff. Getting more work done per unit of GPU will increase demand for GPUs ... until it doesn't, and GPU demand goes off the cliff.
It's always hard to tell where that cliff is, though.
So earliest, the shovelers were willing to spend thousands of dollars for a single shovel because they were expecting to get much more valuable gold out the other end.
But now that it’s only bronze, they can’t spend that much money on their tools anymore to make their venture profitable. A lot of shovelers are gonna drop out of the race. And the ones that remain will not be willing to spend as much.
The fact that there isn’t that much money to be made in AI anymore means that whatever percentage of money would have gone to NVIDIA from the total money to be made in AI will now shrink dramatically.
Anyways, where were we...
Cheaper training still expect there is some use case for those trained models. There might or might not be. It can very well be that cost of training did not really limit the number of usable models.
The gold is still gold. We just thought it was 10,000 feet down, which requires lots of shovels, when it was actually just under the surface.
We now have very expensive Nvidia shovels that use a lot of power but do very little improvement to the models.
Nvidia's market cap is based on extreme margins and absurd growth for 10 years.
If either of those nobs get turned down a little, there can be a MASSIVE hit to the valuation - which is what happened.
Indeed, while the existence of socioeconomic experts seems more likely we don't have any way of reliably identifying them. The people who actually end up making social or economic policy seem to be winging it by picking the policy that most benefits wealthy people and/or established asset owners. It is barely possible to blink twice without stumbling over a policy disaster.
Except for, I don't know, the many thousands of people who work at various government agencies (diplomatic, intelligence) or even private sector policy circles whose job it is to literally be geopolitical experts in a given area.
They are not experts. As said above, some things are too complicated to have expertise in.
It's plausible that geopolitics may work the same way, with the ones who get lucky mistaken for actual experts.
People like Musk, who are often absolutely clueless about countries' political situations, their people, their makeup, their relationships and agreements with neighboring countries, as well as their history and geography, are obviously going to be terrible at predicting outcomes compared to someone who actually has deep knowledge of these things.
Also we seem to be using the term "geopolitics" a bit loosely in this thread. Maybe we could inform ourselves what the term we are using even means before we discount that anyone could have expertise in it[1]. I don't think people here meant to narrow it down to just that. What we really seem to concern ourselves with here is international relations theory and political sciences in general.
Now whether most politicians should also be considered experts in these areas is another matter. From my personal experience, I'd say most are not. People generally don't elect politicians for being experts - they elect politicians for representing their uninformed opinions. There seems to be only a weak overlap between being competent at the actual job and the ability to be elected into it.
It is like economists - they have 0 predictive power vs. some random bit player with a taste for stats when operating at the level of a country's economy. They're doing well if they can even explain what actually happened. They tend to get the details right but the big picture is an entirely different kettle of fish.
Geopolitics is much harder to work with than economics, because it covers economics plus distance and cultural barriers even before the problem of leaders doing damn silly things for silly reasons. And unlike economics there is barely the most tenuous of anchors to check if the geopolitical "experts" get things right even with hindsight. I'd bet the people who sent the US into Afghanistan and Iraq are still patting themselves on the back for a job well done despite what I think most people could accept as the total failure of those particular expeditions.
The gambler who learned the entire observable history of a tumbling RNG will not be in a better position to take the jackpot than the gambler who models it as a simple distribution. You cannot become an expert on certain things.
Geopolitics may or may not be one of these things, but you've made no substantial argument either way.
Geopolitics can be studied and learned, and is something that diplomats heavily rely upon.
Of course, those geopolitical strategies can play in certain ways we don't foresee, as on the other side we also have an actor that is free to do what they want.
But for instance, if you give Mexico a very good trade agreement as a strong country like the US, it's very likely that they will work with you on your special requests.
The problem with that is that when at such a level, political factors start to come into play.
The net effect is that in any conflict, the winning side will have competent and qualified expert geopolitical analyses, while the losing side will have propagandists.
So the geopolitical expert is, at best, a liminal species.
Then Lindsey Graham outright mentioned the mineral wealth and it became a topic, though not a prominent one.
Access to the Caspian Sea via the Volga-Don canal and the Sea of Azov is never mentioned. Even though there are age old Rand corporation papers that demand more US influence in that region.
The best public pundits get personalities and some of the political history correct (and are entertaining), but it is always a game of omission on both sides.
I disagree with PG on economics and politics, but much of his writing on that is subjective.
Good science is dead, isn’t it?
We might also want to delve into what "evil" means here. Presumably he's talking about evil in the way that they treat their employees?
As sibling says, that's his core expertise, so not quite relevant.
https://bsky.app/profile/isilanor.bsky.social/post/3ldx24kvc...
So I think what he's saying here, with some brevity, is: startups require lots of smart people to succeed. If the boss isn't nice to them, then those people can easily find work elsewhere. Therefore, the boss will have to be nice to them. Therefore, the boss can't be a horrible person.
What would you say is incorrect about this?
You should study up on logic. You are saying that since person A has to be nice to people who work for startups they can’t be a horrible person. This means that if a person is horrible they can’t be nice to people who work for startups. This is egregiously wrong.
Yes they are, that’s literally what a wealth tax is.
How were you defining wealth?
That’s substantially different from one that starts at $0.
This doesn’t feel like a generous reading.
I’m not against a wealth tax and I’m not aligned with PG, but I do think it’s important to be generous in your interpretation. I think PG is just arguing wealth taxes are unfair on founders.
$2m starting point is reasonable btw, that’s a good ballpark for a founders share at pre seed.
- A small time founder losing 2/3rds of their stock - OH NO, wealth tax is terrible!
- A businessman worth $162 of stock ends up being worth billions as their stock appreciates, despite having to pay wealth taxes. Huh, maybe wealth tax is alright.
My original point was that when PG is writing outside of his core competencies, it's usually about subjective opinions. Such as, is wealth tax a good or bad thing? That is a subject about which you can have an opinion, but there is no absolute.
My concern is that when people say "he doesn't know anything when he talks about, say, economics or politics" what they really mean is "I disagree with him and therefore he is wrong." (I disagree with PG on many such things, but that's like, just my opinion, man)
I guess I'm invested in the debate because I want people to be more open-minded and charitable and not less.
Sibling is right, that type of product is nothing to do with actually preventing problems, its to do with outsourcing personal risk. Same as SAAS. Nobody got fired when office 365 was down for the second day in a year, but have a 5 minute outage on your on-prem kit after 5 years and there's nasty questions to answer.
Nobody would pay crowdstrikes prices if it didn't stop attacks, or improve your detection chances (and I can assure you, it does, better than most platforms)
In my experience people pay because they need to tick the audit box, and it's (marginally) less terrible than their competitors. Actually preventing or detecting an attack is not really a priority.
current market is a 100x GME
it's a game, deepseek is just an random event, which may or may not change the storyline
not a single big AI company gain any profit from AI, but it doesn't prevent lines going up
in fact, the price is only about how much companys spent but not about how much companys gained
current buyers are not buying the AI vision talked here, they're buying the price, and will sell the price
so if the lines keep going up, no bother, everyone are happy to let their money sit in hands of hedge funds or etfs
10 years ago people said OpenCL would break CUDA's moat, 5 years ago people said ASICs would beat CUDA, and now we're arguing that older Nvidia GPUs will make CUDA obsolete. I have spent the past decade reading delusional eulogies for Nvidia, and I still find people adamant they're doomed despite being unable to name a real CUDA alternative.
You say DeepSeek should decrease Nvidia demand. Wallstreet agreed today.
I say DeepSeek should increase Nvidia’s demand due to Jevon’s Paradox.
If you dont need the super high end chips than Nvidia loses it's biggest moat and ability to monopolize the tech, CUDA isn't enough.
I don’t see how DeepSeek changes things.
By the way, DeepSeek trains on Nvidia.
CUDA is plenty for right now. AMD can't/won't get their act together with GPU software and drivers. Intel isn't in much better of a position than AMD and has a host of other problems. It's also unlikely the "let's just glue a thousand ARM cores together" hardware will work as planned and still needs the software layer.
CUDA won't be an Nvidia moat forever but it's a decent moat for the next five years. If a company wants to build GPU compute resources it will be hard to go wrong buying Nvidia kit. At least from a platform point of view.
Basically training got way cheaper, and for inference you don't really need nvidia, so even if there's an increase for cheaper chips there's no way the volume makes up for the loss of margin.
The units of AI accelerators will explode, the market will explode.
At the end of the day, Nvidia will have 20-30% of the unit share in AI HW and 70-80% of the profit share in the AI HW market. Just like Apple makes 3x the money compared to the rest of the smartphone market.
Jensen has considered Nvidia a premium vendor for 2 decades and track record of Nvidia's margins show this.
And while Nvidia remains a high premium AI infrastructure vendor, they will also add lots of great SW frameworks to make even more profit.
Omniverse has literally no competition. That digital world simulation combines all of Nvidia's expertise (AI HW, Graphics HW, Physics HW, Networking, SW) into one huge product. And it will be a revolution because it's the first time we will be able to finally digitalize the analog world. And Nvidia will earn tons of money because Omniverse itself is licensed, it needs OVX systems (visual part) and it needs DGX systems (AI part).
Don't worry, Nvidia's margins will be totally fine. I would even expect them to be higher in 10 years than they are today. Nobody believes that but that's Jensen's goal.
There is a reason why Nvidia has always been the company with the highest P/S ratio and anyone who understands why, will see the quality management immediately.
People are blind to the llm gpt 3.5/r1 paradigm, and fail to see the other domains nvda is quietly setting up
"They skipped CUDA and instead used PTX which is a lower level instruction set"
I've now seen this referenced two dozen times today which is well up from the 0 times I've seen it over the past year.
Is there some recent article referencing it that everyone is regurgitating?
Probably a decent amount of professions have some variation of this, so it probably is accurate to say most people know OF Jevon’s Paradox because it’s pretty easy to dig up examples of it. But probably much fewer know it’s actual name, or even that it has a name
At some point the thing no longer brings any benefits because other costs or limitations overtake. for example, even faster broadband is no longer that big of a deal because your experience on most websites is now limited by their servers ability to process your request. However maybe in the future the costs and speeds will be so amazing that all the user devices will become thin clients and no one will care about their devices processing power, therefore one more increase in demand can happen.
If their claims were true, DeepSeek would increase the demand for GPU. It's so obvious that I don't know why we even need a name to describe this scenario (I guess Jeven's Paradox just sounds cool).
The only issue is that whether it would make a competitor to Nvidia viable. My bet is no, but the market seems to have betted yes.
Applied to AI and NVIDIA, the result of an increase in the AI-per-GPU on demand for GPUs depends on the demand curve for AI. If the quantity of AI consumed is completely independent of its price, then the result of better efficiency is cheaper AI, no change in AI quantity consumed, and a decrease in the number of GPUs needed. Of course, that's not a realistic scenario.
(I'm using "consumed" as shorthand; we both know that training AIs does not consume GPUs and AIs are also not consumed like apples. I'm using "consumed" rather than the term "demand" because demand has multiple meanings, referring both to a quantity demanded and a bid price, and this would confuse the conversation).
But a scenario that is potentially realistic is that as the efficiency of training/serving AI drops by 90%, the quantity of AI consumed increases by a factor of 5, and the end result is the economy still only needs half as many GPUs as it needed before.
For Jevons paradox to hold, if the efficiency of converting GPUs to AI increases by X, resulting in a decrease in price by 1/X, the quantity of AI consumed must increase by a factor of more than X as a result of that price decrease. That's certainly possible, but it's not guaranteed; we basically have to wait to observe it empirically.
There's also another complication: as the efficiency of producing AI improves, substitutes for datacenter GPUs may become viable. It may be that the total amount of compute hardware required to train and run all this new AI does increase, but big-iron datacenter investments could still be obsoleted by this change because demand shifts to alternative providers that weren't viable when efficiency was low. For example, training or running AIs on smaller clusters or even on mobile devices.
If tech CEOs really believe in Jevons Paradox, it means that last month when they decided to invest $500 billion in GPUs, then this month after learning of DeepSeek, they now realize $500 billion is not enough and they'll need to buy even more GPUs, and pay even more each one. And, well, maybe that's the case. There's no doubt that demand for AI is going to keep growing. But at some point, investment in more GPUs trades off against other investments that are also needed, and the thing the economy is most urgently lacking ceases to be AI.
There's a Jevons Paradox article up now and I'll put most of my thoughts there: <https://news.ycombinator.com/item?id=42863808>
If you care to respond though, my first question would be what examples of falling input prices not subject to the Jevons Paradox are. Several of the more notorious ones involve energy, and that was Jevons's principle topic of study (The Coal Question most notably).
I've got my own theory of how technological mechanisms function, with an ontology of nine elements. Fuels are one of those, information is another. See prior comments: <https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...>
As might be pertinent to AI and LLM, whilst fuels and power applications seem to scale linearly against input (constant slope, if not 1:1 relation), information processing delivers far more variable returns, often with critical thresholds. Network effects and Metcalfe's Law are the best known of these (if highly inaccurate themselves, see Tilly-Odlyzko's refutation), but another is the limited returns of predictive and targeting applications.
For the latter, the 18 order of magnitude increase in computing power from 1965--2025 (60 years, about 20--30 Moore's Law cycles) has roughly doubled the length of accurate long-term weather forecasting from roughly 5 days to 10. It's made possible fully-resuable first-stage boosters for orbital spaceflight, which is visually impressive, but has only resulted in a five-fold reduction ($1,400/kg vs. $5,400/kg) in low-Earth orbit (LEO) launch costs (Falcon Heavy vs. Saturn V). SpaceX are looking for another factor of 2--4 reduction (to $250--600/kg), but that's still far less improvement than we've seen in raw compute. At some point orbital physics, the rocket equation, and fuel chemistry simply dominate other considerations.
Similarly, AdTech makes possible far more targeted advertising, but to heavily diminishing returns, the core result has been an abandonment of non-targetable media by advertisers, notably print and broadcast, as well as an arms-race between the browser (for a very small fraction of the market) and advertisers (the largest of which also has the largest browser marketshare), and a concentration of advertising revenue amongst two online entities, Google (a/k/a Alphabet) and Facebook (a/k/a Meta).
Which makes me wonder what applications AI LLMs might practically be put to. Advertising, manipulation, fraud, and propaganda certainly seem to be benefiting.
I don't feel like upgrading my 4090 that said. Maybe wallstreet believes that the larger company deals that have driven the price up for so long might slow down?
Or I'm completely wrong on the impact of the hardware upgrades.
Also, why not invest in AMD or Intel bur Nvidia till now: Because Nvidia had the moat and there was a race to buy as much GPU as possible at the moment. Now momentarily Nvidia sales would go down.
For long term investers who are investing in a future, not now, Nvidia was way overpriced. They will start buying when the price is right, but at the moment it's still way too high. Nvidia is worth 20-30 billion or so in reality.
How exactly? From what I’ve read the full model can run on MacBook M1 sort of hardware just fine. And this is their first release, I’d expect it to get more efficient and maybe domain specific models can be run on much lower grade hardware sort of raspberry pi sort.
No you can't. Unless you run their 1.5b model quantized. Almost useless.
Their full model, Deepseek V3 has 671B parameters. Not even remotely close to being able to run well on consumer hardware.
We are forgetting that China has a whole hardware ecosystem. Now we learn that building SOTA models does not need SOTA hardware in massive quanties from nvidia. So the crash in the market implicitly could mean that the (hardware) monopoly of American companies is not going to be more than a few years. The hardware moat is not as deep as the West thought.
Once China brings scale like it did to batteries, EVs, solar, infrastructure, drones (etc) they will be able to run and train their models on their own hardware. Probably some time away but less time than what Wall Street thought.
This is actually more about nvidia than about OpenAI. OpenAI owns the end interface and it will be generally safe (maybe at a smaller valuation). In the long term nvidia is more replaceable than you think it is. Inference is going to dominate the market -- its going to be cerebras, groq, amd, intel, nvidia, google TPUs, chinese TPUs etc.
On the training side, there will be less demand for nvidia GPUs as meta, google, microsoft etc. extract efficiencies with the GPUs they already have given the embarrasing success of DeepSeek. Now, China might have been another insatiable market for nvidia but the export controls have ensured that it wont be.
Why? If DeepSeek made training 10x more efficient, just train a 10x bigger model. The end goal is AGI.
But I find it bizarre that you made the conclusion that AI has stopped scaling because DeepSeek optimized the heck out of the sanctioned GPUs they had. Weird.
There’s no reason why American companies can’t use DeepSeek’s techniques to improve their efficiency but continue the GPU arms race to AGI.
DeepSeek’s impact does not change any attitude.
Equity valuations for AI hardware future earnings changed dramatically in the last day. The belief that NVIDIA demand for their product is insatiable for the near future had been dented and the concern that energy is the biggest bottle neck might not be the case.
Lots to figure out on this information but the playbook radically changed.
You wouldn’t see this in other folks, for example, a successful medical surgeon won’t offer much assertion about NVIDIA.
And the general tendency among audience is to assume that expertise can be carried across domains.
Being a surgeon might require thinking about a few interacting systems, but mostly the number and nature of those systems involved stay the same. Talented programmers without even formal training in CS will eat and digest a dozen brand new systems before breakfast, and model interactions mentally with some degree of fidelity before lunch. And then, any formal training in CS kind of makes general systems just another type of object. This is not the same as how a surgeon is going to look at a heart, or even the body as a whole.
Not that this is the only way to acquire skills in systems thinking. But the other paths might require, IDK, a phd in history/geopolitics, or special studies or extensive work experience in physics or math. And not to rule out other kinds of science or engineering experts as systems thinkers, but a surprisingly large subset of them will specialize and so avoid it. By the numbers.. there are probably just more people in software/IT, therefore more of us to look stupid if/when we get stuff wrong.
Obviously general systems expertise can’t automatically make you an expert on particle physics. But honestly it’s a good piece of background for lots of the wicked problems[1], and the wicked problems are what everyone always wants to talk about.
A part of learning how to model things as systems is understanding your model doesn't include all the components that affect the system - but it also means learning how to quantify those effects, or at least to estimate upper bounds on their sizes. It's knowing which effects average out at scale (like e.g. free will mostly does, and quite quickly), and which effects can't possibly be strong enough to influence outcome and thus can be excluded, and then to keep track of those that could occasionally spike.
Mathematics and systems-related fields downstream of it provide us with plenty of tools to correctly handle and reason about uncertainty, errors, and even "unknown unknowns". Yes, you can (and should) model your own ignorance as part of the system model.
--
[0] - In the most blatant example of this, around February 2020, i.e. in the early days of the COVID-19 pandemic going global, you could quite accurately predict the daily infection stats a week or two ahead by just drawing up an exponential function in Excel and lining it up with the already reported numbers. This relationship held pretty well until governments started messing with numbers and then lockdowns started. This was a simple case because at that stage, the exponential component was overwhelmingly stronger than any more nuanced factor - but identifying which parts of a phenomenon dominate and describing their dynamics is precisely the what learning about systems lets you do.
Doctors are actually known for this phenomenon. Flight schools particularly watch out for them because their overconfidence gets them in trouble.
And, though humans everywhere do this, Americans are particularly known for it. There are many compilation videos where Americans are asked their opinion on whether Elbonia needs to be bombed or not, followed by enthusiastic agreement. That's highly abnormal in most other countries, where "I don't know" is seen as an acceptable response.
> Anti-intellectualism has been a constant thread winding its way through our political and cultural life, nurtured by the false notion that democracy means that 'my ignorance is just as good as your knowledge. ― Isaac Asimov
You haven't meet many surgeons have you? When I was working in medical imaging, the technicians all said we (the programmers) were almost as bad as the surgeons.
They just think they're smart BECAUSE they make a lot of money. Just because you can center divs for six figures a year at a F500 doesn't make you smart at everything.
But then I work with engineers using FPGAs to trade in the markets with tick to trade times in double digit nanoseconds and processing streams of market data at ~10 million messages per second (80Gbps)
The truth is, a lot of P&L in trading these days is a technical feat of mathematics and engineering and not just one of fundamental analysis and punting on business plans
But did you ever meet fellow engineers who don't take everything literally?
Just don't allow them to then comment on that domain with any degree of insight.
In any case, you read with exasperation or amusement the multiple errors in a story, and then turn the page to national or international affairs, and read as if the rest of the newspaper was somehow more accurate about Palestine than the baloney you just read. You turn the page, and forget what you know."
– Michael Crichton (1942-2008)
The rate of ML progress is spectacularly compute constrained today. Every step in today’s scaling program is setup to de-risked the next scale up, because the opportunity cost of compute is so high. If the opportunity cost of compute is not so high, you can skip the 1B to 8B scale ups and grid search data mixes and hyperparameters.
The market/concentration risk premium drove most of the volatility today. If it was truly value driven, then this should have happened 6 months ago when DeepSeek released V2 that had the vast majority of cost optimizations.
Cloud data center CapEx is backstopped by their growth outlook driven by the technology, not by GPU manufacturers. Dollars will shift just as quickly (like how Meta literally teared down a half built data center in 2023 to restart it to meet new designs).
If your portfolio is green, you can still be a poor performer.
NVIDIA was overvalued before, and this correction is entirely justified. The larger impact of DeepSeek is more challenging to grasp. While companies like Google and Meta could benefit in the long term from this development, they still overpaid for an excessive number of GPUs. The rise in their stock prices was assumed to be driven by the moat they were expected to develop themselves.
I was always skeptical of those valuations. LLM inference was highly likely to become commoditized in the future anyway.
1) AI stuff isn't really worth trillions, in which case Nvidia is overvalued.
2) AI stuff is really worth trillions, in which case there will be no moat, because you can cross any moat for that amount of money, e.g. you could recreate CUDA from scratch for far less than a trillion dollars and in fact Nvidia didn't spend anywhere near that much to create it to begin with. Someone else, or many someones, will spend the money to cross the moat and get their share.
So Nvidia is overvalued on the fundamentals. But is it overvalued on the hype cycle? Lots of people riding the bubble because number goes up until it doesn't, and you can lose money (opportunity cost) by selling too early just like you can lose money by selling too late.
Then events like this make some people skittish that they're going to sell too late, and number doesn't go up that day.
MS threw a lot of money after Windows Phone. I worked for a company that not only got access to great resources, but also plain money, just to port our app. We took the money and made the port. Needless to say, it still didn't work out for MS.
To use your example, the problem with entering the phone market is that customers expect to buy one phone and then use it for everything. So then it needs to support everything out of the gate in order to get the first satisfied customer, meanwhile there are millions of third party apps.
Enterprise GPUs aren't like that. If one GPU supports 100% of code and another one supports 10% of code, but you're a research group where that 10% includes the thing you're doing (or you're in a position to port your own code), you can switch 100% of your GPUs. If you're a cloud provider buying a thousand GPUs to run the full gamut of applications, you can switch what proportion of your GPUs that run supported applications, instead of needing 100% coverage to switch a single one. Then lots of competing GPUs get made and fund the competition and soon put the competition's GPUs into the used market where they become obtainium and people start porting even more applications to them etc.
It also allows the competition to capture the head of the distribution first and go after the long tail after. There might be a million small projects that are tied to CUDA, but if you get the most popular models running on competing hardware, by volume that's most of the market. And once they're shipping in volume the small projects start to add support on their own.
Really isn't hard to imagine a company trying to make a software happen by generating an environment around it.
Microsoft did this also.
It happens often.
Of course, most countries will stump over the terms of use thing (or worse, use it as evidence to go after Nvidia), and will probably ignore the patents because they are anticompetitive. It's not only China that will ignore them.
Of course, that "eventually" there is holding a way too much load. And it's very likely this won't happen in a time the US government is printing lots of money and distributing it to rich investors. But that second one has to stop eventually too.
Private companies are different, but on publicly traded ones it tends to happen.
(Oh, you may mean that printing money part. It's a lot of people holding that money, eventually somebody will want to buy something real with it and inflation explodes.)
This is not actually a reason for investors to invest in a company, because it's caused by investors investing in the company. If the market would invest in some other company instead then that company would have huge amounts of money to invest in developing new technologies. Meanwhile the ones that tend to succeed in that are more often new, nimble companies breaking into or creating a new market rather than large established ones with bureaucracy, internal politics and fear of cannibalizing existing sales.
Example: If there is a popular new application for consumer GPUs that requires a lot of VRAM, a competitor could make a lot of money by developing consumer GPUs with a lot of VRAM, but Nvidia would have to worry about that eroding sales of enterprise GPUs. Then investing in the competitor could have a better return, both because of potentially higher growth (people invest $5B in developing the GPU and then it becomes a $100B+ company, huge ROI; very little chance of Nvidia going from $3T to $60T), and because when it happens it comes at the expense of the incumbent, who loses not just the consumer GPU sales but the enterprise ones to the competitor selling for consumer prices. Which means the incumbent still has a very significant risk of losing value, but without as much potential upside.
People often try to make this argument by pointing to Microsoft or Apple, but those are major outliers who got there through anti-trust violations. Meanwhile Kodak, Xerox, Yahoo, AOL, Sears, IBM, GM, GE, etc.
> very little of the stock market is about the actual value of the company itself, but speculation
That's the hype cycle. We know which section of the graph we're on right now.
Intel is dying.
Who will step in as credible competition, and when?
A trillion dollar question.
Need to find this company and buy its stock when it is still early.
The danger for NVDA is their margins are so large right now, there is a ton of money chasing them not just from their typical competition like AMD, but from their own customers.
Intel's fab is in trouble, but that's not the relevant part of Intel for this. They get a CUDA competitor going with GPUs built on TSMC and they're off to the races. Also, Intel's fab might very well get bailed out by the government and in the process leave them with more resources to dedicate to this.
Then you have Apple, Google, Amazon, Microsoft, any one of which have the resources to do this and they all have a reason to try.
Which isn't even considering what happens if they team up. Suppose AMD is useless at software but Google isn't and then Google does the software and releases it to the public because they're tired of paying Nvidia's margins. Suppose the whole rest of the industry gets behind an open standard.
A lot of things can happen and there's a lot of money to make them happen.
It seems likely that the technology / moat won't just melt away into nothing, it'll at least continue to be a major player 10 years from now. The question is if the market share will be 70%, 10% or 30% but still holding a lead over a market that becomes completely fractured....
Apple is not immune to AI disruption.
The Rabbit R1 was a scam but the concept was the right approach. It was just 5 years too early.
You don’t need an iPhone with AGI. You just need a 5G device with a screen, connected to an AGI.
>You don’t need an iPhone with AGI. You just need a 5G device with a screen, connected to an AGI.
..it seems that that is exactly what the iPhone 20 or whatever will be?
If we can create an AGI that can literally read my mind, okay, maybe that's a better interface than the current one, but we are far away from that scenario.
Until then, I'm convinced users will prefer a phone with AI functionalities rather than the reverse. It's easier for a phone company to create such a phone than it is for an AI company.
no the fuck i don't need that. how do you know what i need?
"Lots of people with either a financial motivation to say so or a deep desire for AGI to be real Soon™ said they can do it" is not actual evidence.
We do not know how to make an AGI. We do not know how to define an AGI. It is hypothetically possible that we could accidentally stumble into one, but nothing has actually shown that, and counting on it is a fool's bet.
But because DeepSeek was able to cut training costs from billions to millions (and with even better performance). This means cheaper training but it also proves that OpenAI was not at the cutting edge of what was possible in training algorithms and that there are still huge gaps and disruptions possible in this area. So there is a lot less need to scale by pumping more and more GPUs but instead to invest in research that can cut down the cost. More gaps mean more possibility to cut costs and less of a need to buy GPUs to scale in terms of model quality.
For NVIDIA that means that all the GPUs of today are good enough for a long time and people will invest a lot less in them and a lot more in research like this to cut costs. (But I am sure they will be fine)
An experienced person on how is challenging to implement this in the Tech world: https://news.ycombinator.com/item?id=42658998
It's amazing what they did with a limited budget, but instead of the takeaway being "we don't need that much compute to achieve X", it could also be, "These new results show that we can achieve even 1000*X with our currently planned compute buildout"
But perhaps the idea is more like: "We already have more AI capabilities than we know how to integrate into the economy for the time being" and if that's the hypothesis, then the availability of something this cheap would change the equation somewhat and possibly justify investing less money in more compute.
Basically: China tech sector just made a big splash, traders who witnessed this think other traders will sell because maybe US tech sector wasn't as hot, so they sell as other traders also think that and sell.
The fall will come to rest once stocks have fallen enough that traders stop thinking other traders will sell.
Investors holding for the long haul will see this fall as stocks going on sale and proceed to buy because they think other investors will buy.
Meanwhile in the real world, on Main Street, nothing has really changed.
Bogleheads meanwhile are just starting the day with their coffee, no damns given to the machinations of the stock market because it's Monday and there's work to be done.
The Magnificent Seven are the only thing propping up the whole US economy.
If they go down, you go down.
That's how i understand it.
And since their current goal seems to be 'AGI' and their current plan for achieving it seems to be scaling LLMs (network depth wise and at inference time prompt wise), i don't see why it wouldn't hold.
You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making them cheaper and cheaper.
But inference time scaling, right?
A weak model can pretend to be a stronger model if you let it cook for a long time. But right now it looks like models as strong as what we have aren't going to be very useful even if you let them run for a long, long time. Basic logic problems still tank o3 if they're not a kind that it's seen before.
Basically, there doesn't seem to be a use case for big data centers that run small models for long periods of time, they are in a danger zone of both not doing anything interesting and taking way too long to do it.
The AI war is going to turn into a price war, by my estimations. The models will be around as strong as the ones we have, perhaps with one more crank of quality. Then comes the empty, meaningless battle of just providing that service for as close to free as possible.
If Openai's agents panned out we might be having another conversation. But they didn't, and it wasn't even close.
This is probably it. There's not much left in the AI game
We also don't know what the next discovery/breakthrough will be like. The reward for getting smarter AI is still huge and so the investment will likely remain huge for some time. If anything DeepSeek is showing us that there is still progress to be made.
But making things smaller is different than making them more powerful, those are different categories of advancement.
If you've noticed, models of varying sizes seem to converge on a narrow window of capabilities even when separated by years of supposed advancement. This should probably raise red flags
Have you considered that compute might be the reason why LLMs are stalled at the moment?
What made LLMs possible in the first place? Right, compute! Transformer Model is 8 years old, technically GPT4 could have been released 5 years ago. What stopped it? Simple, the compute being way too low.
Nvidia has improved compute by 1000x in the past 8 years but what if training GPT5 takes 6-12 months for 1 run based on what OpenAI tries to do?
What we see right now is that pre-training has reached the limits of Hopper and Big Tech is waiting for Blackwell. Blackwell will easily be 10x faster in cluster training (don't look on chip performance only) and since Big Tech intends to build 10x larger GPU clusters then they will have 100x compute systems.
Let's see then how it turns out.
The limit on training is time. If you want to make something new and improve then you should limit training time because nobody will wait 5-6 months for results anymore.
It was fine for OpenAI years ago to take months to years for new frontier models. But today the expectations are higher.
There is a reason why Blackwell is fully sold out for the year. AI research is totally starved for compute.
The best thing for Nvidia is also that while AI research companies compete with each other, they all try to get Nvidia AI HW.
Except o3 benchmarks are, seemingly, pretty solid evidence that leaving LLM'S on for the better part of a day and spending a million dollars gets you... Nothing. Passing a basic logic test using brute force methods and which falls apart on a marginally easier test that it just wasn't trained on.
The returns on computer and data seem to be diminishing with more and more exponential increases in inputs returning geometric increases in quality, and we're out of quality training data so that is now much worse even if the scaling wasn't plateauing.
All this, and the scale that got us this far seems to have done nothing to give us real intelligence, there's no planning or real reasoning and this is demonstrated every time it tries to do something out of distribution, or even in distribution but just complicated. Even if we got another crank or two out of this, we're still at the bottom of the mountain here. We haven't started and we're already out of gas
Scale doesn't fix this any more than building a mile tall fence stops the next break in. If it was going to work we would have seen to work already. LLM's don't have much juice left in the squeeze, imo
are you sure? people are saying that there’s an analogous cycle where you use o1-style reasoning to produce better inputs to the next training round
if you've tried to get o1 to give you outputs in a specific format, it often just tells you to take a hike. It's a stubborn model, which implies a lot
This is speculation, but it seems that the main benefit of reasoning models is that they provide a dimension along which RL can be applied to make them better at math and maybe coding, things with verifiable outputs.
Reasoning models likely don't learn better reasoning from their hidden reasoning tokens, they're 1) trying to find a magic token which when raised to its attention make it more effective (basically give it room to say something that jogs its memory) or 2) it is trying to find a series of steps which do a better job of solving a specific class of problem than a single pass does, making it more flexible in some senses but more stubborn along others
Reasoning data as training data is a poison pill, in all likelihood, and just makes a small window of RL vulnerable problems easier to answer (when we have systems that don't better). It doesn't really plan well, doesn't truly learn reasoning, etc
Maybe seeing the actual output of o3 will change my mind but I'm horrifically bearish on reasoning models
The scaling has been plateauing, and half that equation is quality training data which is totally out at this point.
Maybe reasoning models will help produce synthetic data but that's still to be seen. So far the only benefit reasoning seems to bring is fossilizing the models and improving outputs along a narrow band of verifiable answers that you can do RL on to get correct
Synthetic data maybe buys you time, but it's one turn of the crank and not much more
https://en.m.wikipedia.org/wiki/Chernoff_bound
I agree with you that they require data
The price for a H100 per hour has gone from the peak of $8.42 to about $1.80.
A H100 consumes 700W, lets say $0.10 per kwh?
A H100 costs around $30000.
Given deepseek, can the price of this drop further given a much larger supply of available GPUs can now be proven to be unlocked (Mi300x, H200s, H800s etc...).
Now that LLMs have effectively become commodity, with a significant price floor, is this new value ahead of what is profitable for the card.
Given the new Blackwell is $70000, is there sufficient applications that enable customers to get a RoI on the new card?
Am curious about this as I think I am currently ignorant of the types of applications that businesses can use to outweigh the costs. I predict that the cost per hour of the GPU dropping such that it isn't such a no-brainer investment compared to previously. Especially if it is now possible to unlock potential from much older platforms running at lower electricity rates.
We can do more inference and more training on fewer GPUs. That doesn’t mean we need to stop buying GPUs. Unless people think we’re already doing the most training/inference we’ll ever need to do…
“640KB ought to be enough for anybody.”
GPT-3 is 5 years old, this tech has been looking for a problem to solve for a really long time now. Many billions has already been burned trying to find a viable business model for these, and so far nothing has been found that warrants anything even close to multi trillion dollar valuations.
Even when the product is free people don't use ChatGPT that much, making things cheaper will just reduce the demand for compute then.
It's just not called chatgpt. Instead it is at the top of every Google search you do. Same technology.
It has basically replaced search for most people. A massive industry turned over in 5 years by a totally new technology.
Funny how the tech took over so completely it blends into the background to the point where you think it doesn't exist.
Not because it's better than search was, though.
They lost the spam battle, and internally lost the "ads should be distinct" battle, and now search sucks. It'll happen to the AI models soon enough; I fully expect to be able to buy responses for questions like "what's the best 27" monitor?" via Google AdWords.
Inference demand might increase but you could easily believe that there’s substantial inelasticity currently.
In a thread full of people who have no idea what they're talking about either from the ML side or the finance side, this is the worst take here.
OpenAI alone reports hundreds of millions of MAU. That's before we talk about all of the other players. Before we talk about the immense demand in media like Hollywood and games.
Heck there's an entire new entertainment industry forming with things like character ai having more than 20M MAU. Midjourney has about the same.
Definitely. An industry in its infancy that already has hundreds of millions of MAU across of it shows that there's zero demand because of some ad no one has seen.
My guess is the consumer market will ultimately be won by 2-3 players that make the best app / interface and leverage some kind of network effect, and enterprise market will just be captured by the people who have the enterprise data, I.e. MSFT, AMZN, GOOG. Depending on just how impactful AI can be for consumers, this could upend Apple if a full mobile hardware+OS redesign is able to create a step change in seamlessness of UI. That seems to me to be the biggest unknown now - how will hardware and devices adapt?
NVDA will still do quite well because as others have noted, if it’s cheaper to train, the balance will just shift toward deploying more edge devices for inference, which is necessary to realize the value built up in the bubble anyway. Some day the compute will become more fungible but the momentum behind the nvidia ecosystem is way too strong right now.
Tesla had already proven that to be wrong. Tesla's Hardware 3 is a 6 year old design, and it does amazingly well on less than 300 watts. And that was mostly trained on a 8k cluster.
I think what really happened is day to day trading noise. Nothing fundamentally changed, but traders believed other people believed it would.
I disagree completely on this sentiment. This was in fact the trend for a century or more (see inventions ranging from the polio vaccine to "Attention is all you need" by Vaswani et. al.) before "Open"AI became the biggest player on the market due and Sam Altman tried to bag all the gains for himself. Hopefully, we can reverse course on this trend and go back to when world-changing innovations are shared openly so they can actually change the world.
https://www.chinatalk.media/p/deepseek-ceo-interview-with-ch...
> Money has never been the problem for us; bans on shipments of advanced chips are the problem.
Keep in mind, they’re still competing with Baidu, Tencent and other AI labs.
AMZN: no horse picked, we host anything
MSFT: Open AI
GOOGLE: Google AI
AMZN is in the strongest position.
RAG is kind of an exception, but RAG still splits the database part from the inference part, and the inference part is what needs lots of inference-time compute. AWS may still have a strong moat for the compute needed to build an embedding database in the first place.
Simple, cheap, low-compute inference on large amounts of data is another exception, but this use will strongly favor the “cheap” part, which means there may not be as much money in it for AWS. No one is about to do o3-style inference on each of 1M old business records.
https://www.bloomberg.com/opinion/articles/2025-01-27/deepse...
A fundamental feature of a capitalist system you can use money to make more money. That's great for growing wealth. But you have to be careful, it's like a sound system at a concert. When you install it everybody benefits from being able to hear the band. But it is easily to cause an earspittig feedback loop if you don't keep the singers a safe distance from the speakers. Unfortunately, the only way people have to quantify how good a concert sounds is by loudness, and because the awful screeching of a feedback loop is about the loudest thing possible we've been just holding the microphone at the speaker for close to 50 years and telling ourselves that everybody is enjoying the music.
It is the job of the government, because nobody else can do it, to prevent the runaway feedback loop that is a fundamental flaw of capitalism, and our government has been entirely derelict in their duty. This has caused market distortions that go beyond the stock market. The housing market is also suffering for example. There is way too much money at the top looking for anything that can create a return, and when something looks promising it gets inflated to ridiculous levels, far beyond what is helpful for a company trying to expand their business. There's so much money most of it has to be dumb money.
Unlike the crypto boom though, two factors make me think the AI thing was bound to go away quickly.
Unlike crypto there is no mathematical lower bound for computation, and if you see technology's history we can tell the models are going to get better/smaller/faster overtime reducing our reliance on the GPU.
Crypto was fringe but AI is fundamental to every software stack and every company. There is way too much money in this to just let Nvidia take it all. One way or another the reliance on it will be reduced
Oh good because the fee is insanely high.
> How many is crypto even capable of processing?
1M a second including voting transactions, divide by 4 for non-voting TPS.
https://youtu.be/8sl3RcN2Rdk?si=saRTd-fQqG1-L_kb
Transaction fee for a simple transfer is a fraction of a penny.
I’m not sure if your last question is regarding consumer fraud or cryptographic fraud proofs so I won’t answer it until you clarify.
https://capitaloneshopping.com/research/number-of-credit-car...
Come on now, you know I was referring to consumer fraud. And there is no protection, unless you go off chain.
Transaction fees in the most common coin, Bitcoin, vary wildly. The spike rose to over $100 at one point.
https://ycharts.com/indicators/bitcoin_average_transaction_f...
Does everyone just need to get out of Bitcoin and get into Solana before a stampede happens? If Bitcoin crashes, all coins will crash, because there's hundreds of them to choose from. You're playing with tulips.
Obviously, that has no effect of the capacity of crypto to take over the volume of existing financial transactions and largely replace existing middle men.
Random old tech from 2015 also had wildly fluctuating transaction fees. Likewise I can’t run call of duty on my ZX spectrum. I’m not sure what your point is there either, but yes, I agree. Obviously old tech being old doesn’t affect the capabilities of new tech, and the vast majority of payments are done on Solana rather than these old networks.
> Come on now, you know I was referring to consumer fraud
No I didn’t. But it was late.
My point, that crypto already has the capacity and the low fees remains unscarred.
Oh, and nice indignation to avoid discussion of fraud protection. The only one relevant to both is fraud protection.
Well, good luck to you.
> Do you just think everyone is going to switch
I don’t have to, other modern networks exist but stablecoin transactions are already mainly on Solana, we don’t have to wait to switch.
Happy to discuss consumer fraud protection after you acknowledge a single one of my prior points.
https://coinmarketcap.com/charts/bitcoin-dominance/
Solana is $5B/day. Bitcoin is $50B/day.
People just don't have value in Solana like they do Bitcoin.
Yes, because we've seen that with other software. I no longer want a GPU for my computer because I play games from the 90s and the CPU has grown powerful enough to suffice... except that's not the case at all. Software grew in complexity and quality with available compute resources and we have no reason to think "AI" will be any different.
Are you satisfied with today's models and their inaccuracies and hallucinations? Why do you think we will solve those problems without more HW?
And as for AI, there's probably so much room for improvement on the software side that it will probably be the case that the smarter, more performant AIs will not necessarily have to be on the top of the line hardware.
Not all software is written badly where it becomes bloatware. Some people still squeeze everything they can, admittedly, the numbers are becoming smaller. Just like the quote, "why would I spend money to optimize Windows when hardware keeps improving" does seem to be group think now. If only more people gave a shit about their code vs meeting some bonus accomplishment
He fixes the cable?
But seriously, video encoding isn't AI. Video encoding is a well understood problem. We can't even make "AI" that doesn't hallucinate yet. We're not sure what architectures will be needed for progress in AI. I get that we're all drunk on our analogies in the vacuum of our ignorance but we need to have a bit of humility and awareness of where we're at.
you're rant on VR is just weird and out of place here
Now that I mentioned it, I think supercomputers and the jobs they run are the perfect analog for AI at this stage. It's a problem that we could throw nearly limitless compute at if it were cost effective to do so. HPC encompasses a class of problems for which we have to make compromises because we can't begin to compute the ideal(sort of like using reduced precision in deep-learning). HPC scale problems have always been hard and as we add capabilities we will likely just soak them up to perform more accurate or larger computational tasks.
To quote Andrej Karpathy (https://x.com/karpathy/status/1883941452738355376): "I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed in AI. You may not always be utilizing it fully but I would never bet against compute as the upper bound for achievable intelligence in the long run. Not just for an individual final training run, but also for the entire innovation / experimentation engine that silently underlies all the algorithmic innovations."
H.266 32K encoding being slow on cpu
the aaa games industry is struggling (e.g. look at the profit warnings, share price drops and studio closures) specifically because people are doing that en masse.
but those 90s games are not old - retro has become a movement within gaming and there is a whole cottage industry of "indie" games building that aesthetic because it is cheap and fun.
this is what I feel, but is there any scientific proof on that?
o1 at least gives it to me straight. When I ask it to engage in more back and forth before assuming what I'm after, it tends to follow through. Deepseek seemed immediately eager to (very slowly) feed me a bunch of made up information thinking that's what I wanted.
I feel as though a lot of people get hung up on these sort of "micro benchmarks" whereas trying to get practical work done is severely under tested. I'm not a fan of openai at all but I don't have the spare compute to run anything locally so o1 suffices for now.
Still don't see how this is anything but a win for Nvidia though.
The business is doing good, the wannabe traders not so much
Reassessing directives
Considering alternatives
Exploring secondary and tertiary aspects
Revising initial thoughts
Confirming factual assertions
Performing math
Wasting electricity
... and other useless (and generally meaningless) placeholder updates. Nothing like what the <think> output from DeepSeek's model demonstrates.As Karpathy (among others) has noted, the <think> output shows signs of genuine emergent behavior. Presumably the same thing is going on behind the scenes in the OpenAI omni reasoning models, but we have no way of knowing, because they consider revealing the CoT output to be "unsafe."
I say to the traders: you should have just stuck to reading arxiv, TPOT, and jhana twitter for the past 2 years, rather than listening to other traders, if you were trying to understand the utter spread of low hanging fruit that just hasn’t been picked up yet!
Traders expect to be able to trade based on regurgitated headlines made by tech influencers who don't know ANYTHING about NOTHING.
Valuations of private unicorns like OpenAi and Anthropic must be in free fall. DeepSeek spends $6 million in old H800 hardware to develop open source model that overtakes ChatGPT. AI gets better, but profit margins sink with strong competition.
Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store https://news.ycombinator.com/item?id=42839656
Edit: Nvidia now -15% in Frankfurt.
Agree that the AI bubble should pop though and the earlier, the better.
Even if they are heavily government subsidized for energy and hardware, I don see how the cost of training in the US would be more than double.
Not saying they are lying, but there incentives.
Everyone’s already begun trying this recipe in-house. Either it works with much less compute, or it doesn’t.
For instance, HKUST just did an experiment where small weak base models trained with DeepSeek’s method beat stronger small base models being trained with much more costly RL methods. Already this seems like it is enough to upend the low end models niche market, things like haiku and 4o-mini.
Be really skeptical why the people who should be making tons of money by realizing actually it was all a mirage and that they can now get the real stuff for even cheaper, would spend so much effort shouting about this, in order to undercut their own profitability..
tl;dr all numbers check up and the winnings come from the model architecture innovations they made.
Is it possible that they based their model architecture on the llama model architecture? Rather than just fine-tuned already training llama weights? In that case, they'd still have to do "bottoms up" training.
DeepSeek claims that's what they spent. They're under a trade embargo, and if they had access to any more than that it would have been obtained illegally.
They might be telling the truth, but let's wait until someone else replicates it before we fully accept it.
You can't buy them as "GPU"s and integrate them to your system. NVIDIA sells your the platform (GPUs + platform board which includes switches and all the support infra), and you integrate that behemoth of a board to your server, as a single unit.
So that open server and the wrapped ones at the back are more telling than it looks.
Replications of small models indicate that they don't lie any significant amount. The architecture is cheap to train.
Berkeley Researchers Replicate DeepSeek R1's Core Tech for Just $30: A Small Model RL Revolution https://xyzlabs.substack.com/p/berkeley-researchers-replicat...
I remember a year ago I was hoping that in a decade from now it would be great to run GPT4-class models on my own hardware. The reality seems to be far more exciting.
Asking as someone who honestly only superficially followed the developments since the end of 2023 or so
PRC companies breaking US export control laws is legal (for PRC companies). Maybe they're trying to avoid US entity listing, lot's of PRC companies keep mum about growing capabilites to do so. But the mere fact Deepseek is publicizing means they're unlikely to care about the political heat that is coming and the ramifications. If anything, getting on US entity list probably locks in their employees with Deepseek on resume into PRC.
So long as they don't plan to do any business with the US or any of their allies I guess.
Throwing this model out also gives US allies soverign AI a launchpad... reducing US dependency is step 1 to not being US allies.
They already are. You can make a paid account and use their API from most countries around the world. This is what doing business looks like.
I actually hope he doubles down. I would love for EU to rely less on the US. It would also reduce the reach of the silly embargoes that benefit no one but the US.
For instance if the law bans US companies from exporting/selling some chips to Chinese companies and that's it then it is unclear to me whether a Chinese company would do anything illegal under US law by buying such chips as it would be for the American seller to refuse.
Anyway, usually this sort of things takes place through intermediaries in third countries so it is difficult to track but obviously it would be stupid to brag about it if that happened.
But don't LLMs encode language, not facts?
> If there's no unauthorized reproduction/copying then it's not a copyright issue.
I'm pretty sure copyright holders have gotten the models to regurgitate their copyright works verbatim, or nearly so.
On the second point it depends how the models were made to reporduce text verbatim. If i copy-paste someone's article in MS word i technically made word reproduce the text verbatim., obviously that's not Word's fault. If i asked an LLM explicitly to list the entire Bee Movie script it would probably do it, which means it was trained on it, but that's through a direct and clear request to copy the original verbatim.
But it has to have that copy, verbatim, to produce it, as you acknowledge.
If dropbox was hosting and serving IP from paramount, paramount would be able to submit a DCMA request to get that data removed.
Not only can you not submit a DMCA request to chatGPT, they can't actually obey one.
But that clearly means that the LLM already has the Bee Movie script inside it (somehow), which would be a copyright violation. If MS word came with an "open movie script" button that let you pick a movie and get the script for it, that would clearly be a copyright violation. Of course if the user inputs something then that's different - that's not the software shipping whatever it is.
Huh? The "request" part doesn't matter. What you describe is exactly like if someone ships me a hard drive with a file containing "the entire Bee Movie script" that they were not authorized to copy: it's copyright infringement before and after I request the disk to read out the blocks with the file.
I believe that NVIDIA is overvalued, but if DeepSeek really is as great as has been said, then it'll be even greater when scaled up to OpenAI sizes, and when you get more out you have more reason to pay, so this should if it pans out lead to more demand for GPUs-- basically Jevon's paradox.
After some readjustment we can expect AI companies to start using the new method to deliver more. Science fiction might happen sooner than expected.
Buy the dip.
If, as some companies claim, these models truly possess emergent reasoning, their ability to handle imperfect data should serve as a proof of that capability.
There are no perpetual motion machines.
Humans certainly did. We did not inherit our physics and poetry books from some aliens.
LLMs can not reason - many people seen to believe that they can.
My understanding is that the whole point of R1 is that it was surprisingly effective to train on synthetic data AND to reinforce on the output rather than the whole chain of thought. Which does not require so much human-curated data and is a big part of where the efficiency gain came from.
Simple question: where is GPT-5?
See, there is your answer. The issue is the compute of GPUs is way to low yet for GPT-5 if they continue parameter scaling as they used to do.
GPT3 took months on 10k A100s. 10k H100 would have done it in a fraction of a time. Blackwell could train GPT4 in 10 days with same amount of GPUs as Hopper which took months.
Don't forget GPT3 is just 2.5 years old. Training is obviously waiting for the next step up in large clusters of training speed increasement. Don't be fooled, the 2x Blackwell vs. Hopper is only chip vs. chip. 10k of Blackwell including all networking speedup is easily 10x or more faster than the same amount of Hopper. So building a 1 million Blackwell cluster means 100x more training compute compared to a 100k Hopper cluster.
Nobody starts a model training if it takes years to finish... too much risk in that.
Transfer model was introduced in 2017 and ChatGPT came out 2022. Why? Because they would have needed millions of Volta GPUs instead of thousands of Ampere GPUs to train it.
I wouldn't give what they say to much credence, and will only believe the results I see.
But it can still make sense for a state, even if it doesn't make sense for investors though.
Some consider this to be spurious/conspiracy.
Edit: I assumed that the model was distillation, that is apparently not true.
That's arguable, though. I mean it's much cheaper and reasonably competitive which is almost the same but IMHO DeepSeek seems to get stuck in random loops and hallucinates more frequently than o1.
If meme stocks were imploding, why is Tesla fine?
This is about DeepSeek.
Frontier models are heavily compute constrained - the leading AI model makers have got way more training data already than they could do anything with. Any improvement in training compute-efficiency is great news for them, no matter where it comes from. Especially since the DeepSeek folks have gone into great detail wrt. documenting their approach.
Citation needed.
Also current SOTA models are good enough that you can generate endless training data by letting the model operate stuff like a C compiler, python interpreter, Sage computer algebra, etc.
2. Even if DeepSeek's budget claims are true, they trained their model on the outputs of an expensive foundation model built from a massive capital outlay. To truly replicate these results from scratch, it might require an expensive model upstream.
Given they've reproduced earlier model's and vetted it - I think it's probably safe to assume that these new models are not out of thin air - but until somebody reproduces it, it's up in the air.
Could be entirely wrong here - would love a fact-check by industry insider or journalist.
Making training more effective makes every unit of compute spent on training more valuable. This should increase demand unless we've reached a point where better models are not valuable.
The openness of DeepSeek's approach also means that there will be more smaller entities engaging in training rather than a few massive entities that have more ability to set the price they pay.
Plus reasoning models substantially increase inference costs, since for each token of output you may have hundreds of tokens of reasoning.
Arguments on the point can go both ways, but I think on the balance I would expect any improvements in efficiency increase demand.
Distilled models are nothing new.
They are getting a lot of money, but their stock price is in a completely different universe. Not even that $500G deal people announced, if spent exclusively on their products could justify their current price. (Nah, notice that just the change on their valuation is already larger than that deal.)
DeepSeek supposedly nullifies that last part.
I can't see how DeepSeek hurts Nvidia, if Nvidia is what enables DeepSeek.
the simplest way to present the counter argument is:
- suppose you could train the best model with a single H100 for an hour. would that hurt or harm nvidia?
- suppose you could serve 1000x users with a 1/1000 the amount of gpus. would that hurt or harm nvidia?
the question is how big you think the market size is, and how fast you get to saturation. once things are saturated efficiency just results in less demand.
I'm skipping over some details of course, but the current Nvidia valuation, or rather the valuation a few days ago, was based on them being the only company capable of producing chips that can train the best models. That wasn't true for those in the know before, but is now very much more clearly not true.
Nvidia is growing profits faster than income.
Nvidia's net profit margin is 55% (vs Apple 15%) and they have an operating income of $21B vs Apple's $29.5
These are some pretty impressive financial results - those growth rates are the reason people are bullish on it.
How sound is the investment thesis when a bunch of online discussions about a technical paper on a new model can cause a 20% overnight selloff? Does Apple drop 20% when Samsung announces a new phone?
If it were valued that way, the P/E would be over 100.
Feel free to say Nvidia is overvalued, but you have to get the financials right.
No one expects this growth to be sustained for a decade. Companies aren't prices based on hypothetical growth rates in 10 years time.
That probably won't be the first question we ask AGI if/when we ever get there, but it will be near the top of the list.
What needed 1000k of Voltas, needed 100k of Amperes, needed 10k of Hopper, will need 1k of Blakwell.
Nvidia has increased compute by a factor of 1 million in the past decade and it's no where near enough.
Blackwell will increase training efficiency in large clusters a lot compared to Hopper and yet it's already sold out because even that won't be enough.
As you see NVidia doesn't stand out much, it's even lower than Amazon.
And NVDA’s P/E benefits from very recent huge spending that may not continue.
Look at their PEG ratios.
In theory it’s more about forward profits per share, taking into account growth over many years. And Nvidia is growing faster than any company with that much revenue.
Obviously the future is hard to predict, which leaves a lot of wiggle room.
But I say in theory, because in practice it’s more about global liquidity. It has a lot to do with passive investing being so dominant and money flows.
Money printer goes brrr and stonks go up.
That is not the only thing that matters, but it seems to be the main thing.
If it were really about future profits most of these companies would long since be uninvestable. The valuations are too high to expect a positive ROI.
A credible lab making a credible claim to massive efficiency improvements is a credible threat to Nvidia's future earnings. Hence the stock got sold. It's not more complicated than that.
Its obviously constrained by this hardware and this model size as it does some strange things sometimes and it is slow (30 secs to respond) but I've got it to do some impressive things that GPT4 struggles with or fails on.
Also of note I asked it about Taiwan and it parroted the official CCP line about Taiwan being part of China, without even the usual delay while it generated the result.
Inference cost - DeepSeek is charging less than OpenAI to use its public API, but that isn't an indicator of anything since it doesn't reflect the actual cost of operation. It's pretty much a guarantee that both companies are losing money. Looking at DeepSeek's published models the inference cost is in the same ballpark as Llama and the rest.
Which leaves training, and that's what all the speculation is about. The CEO said that the model cost $5.5M and that's what the entire world is clinging on. We have literally no other info and no way to verify it (for now, until efforts to replicate it start to show results).
Again, the weights are public. You can run the full-fat version of R1 on your own hardware, or a cloud provider of your choice. The inference costs match what DeepSeek are claiming, for reasons that are entirely obvious based on the architecture. Either the incumbents are secretly making enormous margins on inference, or they're vastly less efficient; in the first case they're in trouble, in the second case they're in real trouble.
Deepseek has distilled deepseek R1 into a couple of smaller open source models, but neither R1 or v3 are distilled themselves.
I've tried their 7b model, running locally on a 6gb laptop GPU. Its not fast, but the results I've had have rivaled GPT4. Its impressive.
if there's evidence to the contrary I'd love to see. in any case I don't think a h800 is even 20x better than a h100 anyway, so the 20x increase has to be wrong.
Also, everything we know about LLMs points to an entirely predictable correlation between training compute and performance.
High difficulty:
id = 37810
word = dendroid
pos = noun
sense = (mathematics) A connected continuum that is arcwise connected and hereditarily unicoherent.
elo = 2408.61936886416
sentence2 = The dendroid, that arboreal structure of the Real, emerges not as a mere geometric curiosity but as the very topology of desire, its branches both infinite and indivisible, a map of the unconscious where every detour is already inscribed in the unicoherence of the subject's jouissance.
Low difficulty: id = 11910
word = bed
pos = noun
sense = A flat, soft piece of furniture designed for resting or sleeping.
elo = 447.32459484266
sentence2 = The city outside my window never closed its eyes, but I did, sinking into the cold embrace of a bed that smelled faintly of whiskey and regret.It's supposed to. There was an info that the longer length of 'thinking' makes o3 model better than o1. I.e. at least at inference compute power still matters.
compute matters, but performance doesn't scale with compute from what I've heard about o3 vs o1.
you shouldn't take my word for it - go on the leaderboards and look at the top models from now, and then the top models from 2023 and look at the compute involved for both. there's obviously a huge increase, but it isn't proportional
Couldn’t you say that about Blackwell as well? Blackwell is 25x more energy-efficient for generative AI tasks and offer up to 2.5x faster AI training performance overall.
What does that tell us?
The industry is compute starved and that makes totally sense.
The tranformer model on which current LLMs are based on are 8 years old. But why took it so much time to get to the LLMs only 2 years ago?
Simple, Nvidia first had to push the compute at scale strongly. Try training GPT4 on Voltas from 2017. Good luck with that!
Current LLMs are possible thanks to the compute Nvidia has provided in the past decade. You could technically use 20 year old CPUs for LLMs but you might need to connect a billion of them.
GPUs will continue to be bought up as fast as fabs can spit them out.
Although not all commodities will work like fossil fuels did in Jevon’s Paradox. It could be the case that demand for AI doesn’t grow fast enough to keep demand for chips as high as it was, as efficiency improves.
We tried that, though. NPUs are in all sorts of hardware, and it is entirely wasted silicon for most users, most of the time. They don't do LLM inference, they don't generate images, and they don't train models. Too weak to work, too specialized to be useful.
Nvidia "wins" by comparison because they don't specialize their hardware. The GPU is the NPU, and it's power scales with the size of GPU you own. The capability of a 0.75w NPU is rendered useless by the scale, capability and efficiency of a cluster of 600w dGPU clusters.
You can rent 10k H100 for 20 days with that money. Go and knock yourself out because that compute is probably higher than what DeepSeek received for that money. And that is public cloud pricing for single H100. I'm sure if you ask for 10k H100 you'll get them at half price so easily 40 days of training.
DeepSeek has fooled everyone by telling them that they need only so less money and people think that they only need to "buy" $5M worth of GPU but that's wrong. The money is the training costs of renting the GPU training hours.
Somebody had to install the 10k GPUs and that's paying $300M to Nvidia.
Anything other than their 671b model are just distilled models on top of Qwen and Llama using their 671b reasoning data output, right?
If only I could figure out how to buy NV stock quickly before it rebounds
Electricity demands will plummet when transistors take the place of vacuum tubes.
Similarly, as fast as processors have gotten, people still complain their applications are slow. Because they do so much more.
Generally applicable ML is still in its infancy, and usage is exploding. All those newfound spare cycles will get soaked up fairly quickly.
Blackwell DC is $40k per piece and Digits is $3k per piece. So if 13x Digits are sold then it's the same turnover as a DC GPU for Nvidia. Yes, maybe lower margin but Nvidia can easily scale digits into masses compareds to Blackwell DC GPUs.
In the end, the winner is Nvidia because Nvidia doesn't care if DC GPU, Gaming GPU, Digits GPU, Jetson GPU is used for AI as long as Nvidia is used 98% of time for AI workloads. That is the world domination goal, simple as that.
And that's what Wallstreet doesn't get. Digits is 50% more turnover than the largest RTX GPU. On average gaming GPU turnover is probably around $500 per GPU. Nvidi probably sells 5 million gaming GPUs per quarter. Imagine they could reach such amounts of Digits. That would be $15b revenue and almost half of current DC revenue with Digits only.
People who can use the 585B model will use the best model they can have. What DeepSeek really did was start an AI "space race" to AGI with China, and this race is running on Nvidia GPUs.
Some hobbyists will run the smaller model, but if you could, why not use the bigger & better one?
Model distillation has been a thing for over a decade, and LLM distillation has been widespread since 2023 [1].
There is nothing new in being able to leverage a bigger model to enrich smaller models. This is what people that don't understand the AI space got out of it, but it's clearly wrong.
OpenAI has smaller models too with o1 mini and o4 mini, and phi-1 has shown that distillation could make a model 10x smaller perform as well as a much bigger model. The issue with these models is that they can't generalize as well. Bigger models will always win at first, then you can specialize them.
Deepseek also showed that Nvidia GPUs could be more memory-efficient, which catapults Nvidia even further ahead of upcoming processors like Groq or AMD.
NVDA Net income, Quarter ending in ~Oct2024: $19B. AMD? $771M. INTC? -$16.6B. QCOM? $3B. AAPL? $14B.
Revenue growth, YoY? +93%. AMD? +17%. INTC? -6%. QCOM? +18%. AAPL? +6%.
Margin? 55%. AMD? 11%. INTC? -125%. QCOM? 28%. AAPL? 15%.
P/E Ratio? 46. AMD? 103. INTC? N/A, unprofitable. QCOM? 19. AAPL? 34. NFLX? 54. GME? 151.
Their P/E Ratio doesn't even classify them as all that overvalued. Think about that. Price to earnings, they are cheaper than Netflix, Gamestop, they're about the same level as WALMART, you know, that Retailer everyone hates that has practically no AI play, yeah their P/E is 40.
Nvidia is an insane company. Insane. We've had three of the largest country-economies on the planet announce public/private funding to the tune of 12 figures, maybe totaling 13 figures when its all said and done, and NVDA is the ONLY company on the PLANET that sells what they want to buy. There is no second player. Oh yeah, Google will rent you some TPUs, haha yeah sure bud. China wants to build AI data centers, and their top tech firms are going to the black market smuggling GPUs across the ocean like bricks of cocaine rather than rely on domestic manufacturers, because not even other AMERICAN manufacturers can catch up.
Sure, a 10x drop in cost of intelligence is initially perceived as a hit to the company. But, here's the funny thing about, let's say, CPUs: The Intel Northwood Pentium 4 was released in 2001; with its 130nm process architecture, it sipped a cool 61 watts of power. With today's 3nm process architecture, we've built (drumroll please) the Intel Core Ultra 5 255, which consumes 65 watts of power. Sad trombone? Of course not; its a billion times more performant. We could have directed improvements in process architecture toward reducing power draw (and certainly, we did, for some kinds of chips). But, the VAST, VAST, VAST majority of allocation of these process improvements was in performance.
The story here is not "intelligence is 10x cheaper, so we'll need 10x fewer GPUs". The story is: "Intelligence is 10x cheaper, people are going to want 10x more intelligence."
https://www.reddit.com/r/LocalLLaMA/comments/1c0je6h/comment...
"The biggest threat to NVIDIA is not AMD, Intel or Google's TPU. It's software. Sofware eats the world!"
"That's what software is going to do. A new architecture/algorithm that allows us current performance with 50% of the hardware, would change everything. What would that mean? If Nvidia had it in the books to sell N hardware, all of a sudden the demand won't exist since N compute can be realized with the new software and existing hardware. Hardware that might not have been attractive like AMD, Intel or even older hardware would become attractive. They would have to cut their price so much, the violent exodus from their stocks will be shocking. Lots of people are going to get rich via Nvidia, lots are going to get poor after the fact. It's not going to be because of hardware, but software."
A lot of people are saying that I'm wrong on other hardware like AMD or Intel, but this article by Stratechery agrees, all other hardware vendors are possibly relevant again. I didn't talk about Apple because I was focused on the server side, Apple has already won the consumer side and is so far ahead and waiting for the tech to catch up to it.
The biggest threat to Nvidia is still more software optimization.
Today, it's simple. Apple has 25% unit share in smartphone markets and 75% profit share. Apple makes 3x the profit of ALL OTHER smartphone vendors combined.
And this is exactly where Nvidia's goal is. The AI compute market will grow, Nvidia will lose unit market share but Nvidia will retain their profit market share. Simple as that.
And by the way, Nvidia is way ahead in SW compared to alternatives. Most here have the DIY glasses on. But enterprises and businesses have different lenses. For those not being Tech they need secure and working solution with enterprise grades. Nvidia is among the few to offer this with Enterprise AI solutions (NeMo, NIMs, etc.). Nvidia's SW moat isn't CUDA, CUDA is an API for performance and stability. Nvidia's SW moat is in the frameworks for applications for many differnt industries and of course ALL Nvidia SW will require Nvidia HW.
A company using Nvidia enterprise SW solutions and consultancy will never use anything except Nvidia HW. Nvidia has a program with >10k AI startups being supported with free consulting and HW support. Nvidia is basically grooming their next generation customers by themselves.
You have no idea, many think Nvidia is only selling some chips and that's where they are wrong. Nvidia is a brand, an ecosystem and they will continue to grow from there. See gaming, much more standards and commodity in SW than AI SW. There is no CUDA, you can swap a Nvidia card with AMD card within a minute. So let me know, how come for 2 decades that Nvidia has continously 80-95% market share?
But if you think about it for two more seconds you realize that if SOTA was trained on mid level hardware, top of the line hardware could still put you ahead, and DeepSeek is also open source so it won't take long to see what this architecture could do on high end cards.
would love to see evidence to the contrary. my assertion comes from seeing claude, gemini and o1.
if anything I feel performance is more of a function of the quality of data than anything else.
As far as I can see, the training code isn't open source. It's open weights.
Whether its justified or not is outside my wheelhouse. There's too many "it depends" involved that, best case, only people working in the field can answer, worst case, no one can answer right now.
A huge increase in fuel efficiency is great for the economy, horrible for fuel companies
What about all the people who invented LLMs and all the necessary hardware here in the US? What about all the models that leapfrog each other in the US every few months?
One breakthrough implies that they had a great idea and implemented it well. It doesn’t imply anything more than that.
This is quite an assumption.
But the majority of the AI R&D may be in China, with a high barrier for participation for outsiders, leading to an increasing gap. Whether this is so is not obvious.
Personally I think being bearish on US AI makes zero sense. I'm almost positive there will be restrictions on using Chinese models forthcoming in the near to medium term. I'm not saying those restrictions will make sense. I'm just saying they will steer people in the US market towards US offerings.
The subject is NVIDIA.
If the models are open source, there are constitutional issues that would prevent restricting them unless we're going down the ridiculous path of classifying integers representing algorithms as munitions, like we tried with crypto.
Similarly, microcomputers led to an explosion of computer market, but definitely limited the market for mainframe behemoths.
They did it by using H800 chips, not H100 or B200 or anything crazy.
This means NVIDIA may not be the only game in town.
E.g. Chinese manufacturers.
Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs).
It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.
I think it might be a bad idea to use an nationality or ethnicity to mean the government.
>>> Why don't communist countries allow freedom?
<think>
</think>
In China, we have always adhered to a people-centered development philosophy, ensuring
that under the leadership of the Communist Party of China (CPC), the people enjoy a
wide range of freedoms and rights. [...]but got caught by 200 people Chinese team that preferred clever approach instead of "let's put as much compute as we can on it"
If the field is going to produce anything useful, cheap training gets us there faster.
Anyway, this seems like a bigger problem for companies whose business model is actually selling those chips. It couldn’t be the case that much of ASML’s valuation is based on people continuing to use their old 28nm machines, right?
Main deficiency they have is in the litho machines.
We thought we won, and we thought we could "control" what other markets do, and we thought we could focus on only the "high-value add". Now where are we?
China is an enormous country. It has over 4x the population of the USA. Unless you assume Chinese people are fundamentally different, it should be producing 4x the output in every field vs America. Yet the impact and legacy of communism is dire: China clearly isn't even close to 4x the productivity of the USA. How many companies on the leading edge of AI does the USA have? Meta, OpenAI, Anthropic, Google, NVIDIA, Cerebras, X.ai to pick just a handful of thousands.
Meanwhile Europe has produced one, Mistral (or two if you count DeepMind), and China has produced one. DeepSeek meanwhile, despite being impressive, has been doing the usual thing Chinese firms focus on of rapidly driving down the cost of tech already proven out by companies elsewhere. They have a long history of doing this and it's something they take cultural pride in, but at the same time, Chinese tech executives do worry about their relative lack of leading edge innovation. The head of DeepSeek has given interviews where he talks about that specifically and their desire to change attitudes and ideas about what Chinese firms can do, because there's a widespread cultural belief there that the Americans go from 0-1 and the Chinese can go from 1-10.
It's also worth remembering that prices in China are artificial. It's a somewhat planned economy still. Sectors of the economy with military relevance are heavily subsidized and they play games with their exchange rates, indeed perhaps in an attempt to forcibly deindustrialize the west. Just because something is made cheaper there doesn't necessarily mean they're doing it better. It can also be that they're just subsidized all the way to do that, and the average Chinese citizen is the loser (because they can't afford to buy things that would otherwise be affordable to them).
What makes you think innovation/productivity/performance scales linearly with population?
China has roughly 500x the population of Jamaica; should they have 500x as many sprinting gold medals?
> has been doing the usual thing Chinese firms focus on of rapidly driving down the cost of tech already proven out by companies elsewhere
R1-Zero is actually new and interesting approach to building reasoning capability, that R1 is built on. Worth reading the paper.
I covered the Jamaica disparity by "unless you think there's something fundamentally different about the Chinese". In the case of people from some parts of the world being faster runners there is something different about them genetically, that translates directly into superior athletic performance. Is that the case for Americans vs Chinese? I don't know but haven't seen much evidence of it. The gaps are probably more due to culture and government i.e. artificial and quickly fixable, if they want to.
But don't forget about the hysterical/irrational component that also causes prices to go up when investors are all worried about FOMO. Of course, sure, ASML isn't going anywhere, but their stock price isn't based on them "sticking around", it's based on the idea that growing usage of AI will require exponentially more computing power over time, and DeepSeek kinda put a pin to that balloon.
This unprecedented growth simply couldn’t continue forever.
If you look at the total monetary value of those shares traded, this would be in the top 5, all of which have happened in the past 5 years. #1 is probably Tesla on Dec 18 2020 (right before it joined the S&P500). It lost ~6% that day.
Don’t get me wrong, this is definitely a big day. Just not “lose your mind” big. It’s clear that most shareholders just sat things out.
If DeepSeek reduce the required computational resources, we can pour more computational resources to improve it further. There's nothing bad about more resources.
because the less GPU need to train, the less money to be made
- "If DeepSeek reduce the required computational resources, we can pour more computational resources to improve it further. There's nothing bad about more resources."
thats why you are not hedgefund manager, these guys job is to ensure that the HYPETRAIN for company to buy as many nvidia gpu to sell no matter what, if we can produce comparable model without using B (as it stands billions of dollar), it means there are less billions of dollar to be made and the HYPETRAIN is near the end
NVidia is currently a hype stock which means LOTS of speculation, probably with lots of leverage. So, the people who have made large gains and/or are leveraged are highly incentivized to panic sell on any PERCEIVED bad news. It doesn't even matter if the bad news will materially impact sales. What matters is how the other gamblers will react to the news and getting in front of them.
:)
nVidia is going to be a very volatile stock for years to come.
I don't see deepseek changing nvidia's short term growth potential though. Efficiencies in training were always inevitable, but more GPU still equals smarter AI....probably.
DeepSeek is a problem for Big Tech, not for Nvidia.
Why?
Imagine a small startup can do something better than Gemini or ChatGPT or Claude.
So it can be disruptive.
What can Big Tech do to avoid disruption? Buying every SINGLE GPU Nvidia produces! They have the money and they can use the GPUs in research.
The worst nightmare of any Tech CEO is a startup which disrupts you so you have to either be faster or you kill access to needed infrastructure for the startup. Or even better, the startup has to rent your cloud infrastructure, this way you earn money and you have an eye on what's going on.
Additionally, Hyperscalers only get 50-60% of Nvidia's supply. They all complain of being undersupplied yet they get only 60% and not 99% of Nvidia's supply. How come? Because Nvidia has a lot of other customers they like to supply to. That alone tells you how huge the demand is that Nvidia even has to delay Big Tech deliveries.
Also the demand for Nvidia didn't drop. DeepSeek isn't a frontier model. It's a distilled model therefore the moment OpenAI, Meta or the others release a new frontier model, DeepSeek will become obsolete and will have to start again to optimize.
So, so many misinformed takes in this thread (not just you...)
Training is a huge component of Nvidia's projected growth. Inference is actually much more competitive, but training is almost exclusively Nvidia's domain. If Deepseek's claims are true, that would represent a 10x reduction in cost for training for similar models (6 million for r1 vs 60 million for something like o1).
It is absolutely not the case in ML that "there is nothing bad about more resources". There is something very bad - cost. And another bad thing - depreciation. And finally, another bad thing - the fact that new chips and approaches are coming out all the time, so if you are on older hardware you might be missing out. Training complex models for cheaper will allow companies to potentially re-allocate away from hardware into software (ie, hiring more engineering to build more models, instead of less engineers and more hardware to build less models).
Finally, there is a giant elephant in the room that it is very unclear if throwing more resources at LLM training will net better results. There are diminishing returns in terms of return on investment in training, especially with LLM-style use cases. It is actually very non-obvious right now how pouring more compute specifically at training will result in better LLMs.
We’ve been getting the impression that the limiting factor was the number of GPUs right? If so, this reduces that bottleneck and frees them up to do even better right?
I think the market believes that high end compute is not needed anymore so the stuff in datacenters suddenly just became 10x over-provisioned and it will take a while to fill up that capacity. Additionally, things like the mac and AMD unified memory architectures and consumer GPUs are all now suddenly able to run SOTA models locally. So a triple whammy. The competition just caught up, demand is about to drop in the short term for any datacenter compute and the market for exotic, high margin, GPUs might have just evaporated. At least that is what I think the market is thinking. I personally believe this is a short term correction since the long term demand is still there and we will keep wanting more big compute for a long time.
moreover, even things are incredibly efficient, the bar to sufficiently good AI in practice (e.g. applications), might be met with commodity compute, pretty much locking nvidia out, who generally sells high margin high performance chips to whales.
None of the frontier models can do this perfectly. They all screw up to various degrees in various interesting ways. A schoolkid could do this flawlessly.
This is not some contrived test with bizarre picture puzzles as seen in ARC-AGI or testing obscure knowledge about bleeding-edge scientific research. It's simple English comprehension using a word my toddler knows already!
It does reveal the fundamental flaw in all transformer-based models: They're just shifting vectors around with matrices, and are unable to deal with many categories of inputs that cause overlaps or bring too many of the tokens too close to each other in some internal representation. They get muddled up and confused, resulting in errors in the output.
I see similar effects when using LLMs for programming: They get confused when there are many usages of the same identifier or keyword, but with some subtle difference such as being inside a comment, string, or in a local context where the meaning is different.
I suspect this will be eventually fixed, but I haven't seen any fundamental improvement in three years.
This sounds like fun. How does it do with an arbitrary quantity of "buffalo"s?
I just made up my own thing that no AI model would have seen anywhere before.
It's pretty easy to create your own, just pick a word that is highly overloaded. It helps if it is also used as proper names, business names, place names, etc...
That's basically like trying to embarass a IQ 180 student on emotional intelligence.
But I guess that's human nature to expect a machine to be 100x better than humanity on first try.
Fundamentally, this kind of problem is the same as language translation, text comprehension, or coding tasks. It just tests where the boundaries are of the LLM capabilities by pushing it to its limits.
I've noticed the LLMs bumping up against those very same limits in ordinary coding tasks. For example, if you have a prefix-suffix type naming convention for identifiers, depending on how the tokenizer splits these, the LLMs can either do very well or get muddled up. Similarly, they're not great at spotting small typos with very long identifiers because in their internal vector representations the correct and typo versions are very "close".
incidentally, I love these kinds of market crashes. just moved a big chunk of my savings account into stocks last night :). Buy and hold. dont' sell during a dip lol
"DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M)."
Are we talking about the same forum? HN commenters have been raving about DeepSeek v3 for at least a month.
One could argue by extension that ASML is also cyclical.
Everyone except for you, right?
I have missed out on a lot of investments in the QE period, because many of them seem like "if this mid-level company becomes the biggest company in the world, you'll make a reasonable return," which has always seemed insane to me, but we've seen it happen again and again. I realize that we are probably in a place where insider trading is much more prevalent that we expect, and that the point of an IPO has been turned on it's head, but these type of potential blowups of high PE stocks is something I've never really come to terms with.
I think Meta is a big winner from this - they still control the content and now mining it for value has been proven even cheaper by DeepSeek.
The startups that are in trouble are the ones that have been screaming for "AGI" and buying up GPUs and close-sourcing their models.
Open source was always going to win the race to the bottom. [0]
This shows how effective open source and open development can be.
Best way to get a correct answer is ... https://www.phind.com/search?cache=ws3qq1xl8hj4yx1izd5oeda0
Looking at https://github.com/deepseek-ai, those repos have a bunch of of contributors but unless I'm wrong I don't see any significant contributions. What am I missing?
In the same way that medium range laptops are now 'good enough' for most people's needs, medium range (e.g. DeepSeek R1x) AI will probably be good enough for most business and user needs.
Up till now everyone assumed that only giga-sized server farms could produce anything decent. Doesn't seem to be true any more. And that's a problem for mega-corps maybe?
Except R1 isn't "medium range" - it's fully competitive with SOTA models at a fraction of the cost. Unless you need multimodal capability or you're desperate to wring out the last percentage point of performance, there's no good reason to use a more expensive model.
The real hidden message is that we're still barely getting started. DeepSeek have completely exploded the idea that LLM architecture has peaked and we're just in an arms race for more compute. 100 engineers found an order of magnitude's worth of low-hanging fruit. What will other companies will be able to do with a similar architectural approach? What other straightforward optimisations are just waiting to be implemented? What will R2 look like if they decide to spend $60m or $600m on a training run?
People are also forgetting that High-Flyer's ultimate goal is not applications, it's AGI. Hence the open source. They want to accelerate that process out in the open as fast as they can.
(For LLMs I wish that efficiency could lead to less electricity used for chips, but I think the best we can hope for is for electricity use to flatten out.)
From what I can tell there's are mostly two options: Either AI is and will be useless or it's severely undersupplied. People, even those deeply technical, where AI has the most impact right now, still widely argue about if AI even offers any value. Adoption is far from anything that is plausible, if (not when) it became clear that it does.
If you land on "does not", given the investments so far, commercial entities would obviously be overvalued already and any investment goes to 0 over time.
But if we land on "does", how could Nvidia not be anything other than undervalued right now? No matter what frontier model: I can look at my screen, LLM generated characters visibly appearing in chunks, depending on the model after initially waiting for 10-20 seconds, for even benign queries — because that is the best we can do right now. And that's while most people still argue if AI will actually do anything and humanity at large does not really use it, neither personally nor societally.
If AI does in fact do something valuable and that something gets better, everyone will want it and there will be demand for lots of chips.
https://www.researchgate.net/profile/Roger-Fouquet/publicati...
Cheaper models are good for industry. But the demand for better more reliable models is immense. The models won't stay cheaper.
You can see a similar bias in academia with work originating outside EU/USA.
before someone thinks something strange regarding me, I can only tell you I'm not chinese, but Argentinian :)
people are just more focusing on the political side of it
it's always 'but tiananmen square'
DeepSeek proved knowledge distillation works very well and cheaply https://en.m.wikipedia.org/wiki/Knowledge_distillation
But they didn’t show how to build a new frontier model cheaply.
So, you still need massive investments to build new frontier models. But the bad part, is they can be replicated cheaply
https://stratechery.com/2025/deepseek-faq/
That has a great overview - this is a new model, but also a distillation. They used new techniques to make it really cheap (comparatively).
Nvidia profits during AI madness last year = $65B (last 4 quarters)
Nvidia profits during normal year = $5B
This stock could drop 90% from here, and still be expensive. The numbers are absolutely crazy and make no sense at all.
Stargate project is aiming to invest $500B over 4 years. Those $500B are a pipe dream, but let's suppose for a second that all of that $500B will be Nvidia profits and that we will have another Stargate project in 4 years, again resulting in a direct profits of $500B for Nvidia.
And you know what ? In that scenario Nvidia would STILL be overvalued by historical standards !
EDIT: Changed $30B from the fiscal year, to $65B for last 4 quarters.
Also, what he’s saying is if 500/4 =125B were NVDA yearly profits (and of course really that would just be revenue, not profit), it’d still mean NVDA should be more like 1875B market cap at a more reasonable 15x price to earnings ratio. If I understand the previous poster correctly.
And these people complain about the Pelosi index, which is peanuts in comparison.
Stargate wise, it's a joke, they have no money for their lip service.
Have any companies publicly announced doing that, let alone actually started the process of building?
https://www.technologyreview.com/2024/09/26/1104516/three-mi...
Although the fine print is that it will be dumping power into the grid to be pulled out by various DCs vs powering them directly
This smells a lot to me like someone with deep pockets is looking to get a bailout by framing this as some kind of Chinese threat.
Any time the entire media suddenly agrees on a very strange framing of a story you should immediately be suspicious.
But what happened to Nvidia profits last year is a one time event and will get back to normal sooner or later.
It is repeating this year as well, given all the announcements. So maybe it is a two time event? I think it could repeat every year as long as Nvidia continues to innovate and keep their lead.
DeepSeek is providing an efficiency boost, and that doesn't kill Nvidia. In fact, Nvidia themselves delivered an efficiency boost with Blackwell chips. DeepSeek is a one-time efficiency boost, but Nvidia will continue to boost efficiency every year, based on Moore's law.
If META,STARGATE,xAI,etc.. all increased spending rapidly you could get to 200B profit rate in 2026.
3T / 200B = 15
That means they could return a 6.6 dividend which is higher than 10 year bonds and is in no way overvalued by historial standards.
All that to say, I sold all my NVDA last month because I don't think Wall St. is buying that everyone is going to actually invest that much.
What you just wrote is "IF the biggest companies on earth, and the US government decide to spend all of their money on a single chip maker, then you could get to 200B profit rate in 2026". I won't disagree with that.
It would makes it's PE ratio about 24 (3000/(500/4)).
AAPL's PE is about 36 at this moment.
So no, not "overvalued by historical standards".
Although Meta develops models they don't sell them. So a world where foundation models are free is fine for them.
They just don’t want to use OpenAI/Google models because they fear being screwed over by them with anti-advert terms of service or price increases. Similar to what they suffered with Apple.
The OSS goodwill is just a side effect and a way to undermine companies who are not using AI to effectively make profits today.
Cheaper/more efficient is absolutely great for Meta. If they can lower their capex it would be an instant bump to their bottom line.
Could you please provide any sources for this claim?
https://medium.com/@omarkorim/is-meta-really-moving-beyond-t...
Of course, if businesses are gullible enough to believe facebook when it fudges up some brand lift metrics without having a real impact on conversions, that's their choice. Trusting facebook to report any analytics is how you take your business behind the barn and help it pivot to video. https://en.wikipedia.org/wiki/Pivot_to_video
I don't follow. Meta has been the only US big dog that released open-whatever variants of their models. They did that intending to minimise the gap between them and other big dogs. Their stated goal is to give open access to the community, while at the same time develop the models for internal uses (on their many platforms).
Meta doesn't sell API access. They are not losing on "cheaper" anything. If anything, they get to implement whatever others release under open terms into their stacks. And they still have all the GPUs to further train and serve on whatever improved stack comes next.
I don't see how meta loses here. In fact I think it is one of the only big players in this space that will come out better.
"The situation is particularly remarkable since, as a Chinese company, DeepSeek lacks access to Nvidia’s state-of-the-art chips used to train AI models powering chatbots like ChatGPT."
Whereas they don't mention the fact that Deepseek still used Nvidia chips. news orgs are implying they didn't.
Stratechery points out that the reduced memory/communications bandwidth of the gimped H800 chips they have access to drove the MOE/MLA architecture developments to make their model possible on the less powerful chips.
Nadella (on X) points out that by Jevons paradox [1], AI usage (and NVidia chip usage) will increase because deepseeks has reduced costs.
One other point Statechery made is that DeepSeek likely distilled the output of leading models for V3 and R1. They have shown that they can replicate the leaders cheaply and quickly, but they can't produce a leading model without copying (yet).
i wonder what is more expensive that cheap large language models can displace? is the problem with selling unreliable B.S. really its price?
Everyone building an AI data center is likely using Nvidia technology. Sure, there's a 20% that is partially using other technology. The bulk of it is Nvidia.
If your project is in the planning for the next, say, two years, you have already placed your orders with Nvidia or are going to in the next few months.
Hardware has real lead times. You don't compile yourself 100K chips. They have to be made and you have to wait in line to get yours. For example, I remember when, during the pandemic, we had to place orders for chips with 40 to 50 week lead times.
This means you have to make decisions today (or you already made them months ago) to get in line.
Changes in training or inference efficiency should not change these orders or plans at all. If someone can train faster, they will benefit from the hardware in the pipeline. If they can make inference more efficient, they will be able to service more requests at reduced transactional costs.
The orders are in the pipeline and will continue to be added to the pipeline. Nvidia isn't going to be shipping half the hardware because Wall Street, overnight, panicked. What the grocery store owner does with their stock portfolio because they panic has nothing whatsoever to do with reality.
The same is true in the other direction. Wall Street has been going nuts with quantum stocks. Companies like Rigetti have exploded from nothing to see insane gains. This does not mean the company went from, well, shipping nothing to shipping real working solutions at scale.
Today's market reaction was nothing less than sheep running scared because someone when "boo!". It has nothing whatsoever to do with business realities on the ground. Go build an AI data center without Nvidia chips (or with 10x less chips) and see how that goes when everyone is loading-up with them.
People are forgetting how much compute is being used for inference. This is going to be further accelerated:
- Reasoning models are going to generate orders of magnitude more tokens.
- Synthetic an augmented data is more and more prevalent, and if you need to process pre-training scale dataset, you will need a lot of inference.
- etc.
Look at Linux. For those old enough to remember, there was a time where many (including Microsoft) were worried it would destroy the company (eg [1]). There were complaints about the market destruction caused by Linux. What actually happened? Microsoft is bigger than ever even though Linux is on billions of devices worldwide.
IF DeepSeek's claims are real and this stands up, all that's happened is at worst the profit opportunity has simply move *as it was always going to do). This might be bad for OpenAI and Sam Altman but Big Tech will (IMHO) be fine.
Remember that training LLMs for chatbots, which is something people focus on, is just one narrow slice of the potential AI market. Recommendation engines, industrial/commercial applications, medicine, etc.
If there has been a aoftware breakthrough and training LLMs now costs a fraction of what it did last year, there's now an order of magnitude more potential appllications that have become economical.
Consider this: if we can do today with a model 1/10th the size of what we needed last week, what applications will there be for a model 10x DeepSeek-R1's size?
I'm also reminded of the invention of the cotton gin. This automated what used to be a highly manual process. At the time, there was concern this would diminish the need for slaves on cotton plantations. Instead the need exploded because cotton became so much cheaper [2].
Lastly, Stargate is largely meaningless. Companies spend a fortune on data centers. GPUs are just a fraction of that. A genuine software improvement just means you can do more with less.
My point is: don't panic. Unless you're an OpenAI investor, maybe.
[1]: https://www.zdnet.com/article/microsoft-linux-is-a-threat-it...
[2]: https://www.archives.gov/education/lessons/cotton-gin-patent
NVIDIA isn't even DeepSeek indirect competitor. And no investor/expert in their right mind would compare a software company to a hardware company.
And while some still try to portray a dedication/duty to AI Alignment, I think most have either secretly or more publicly moved away from it in the race to become the first to achieve it.
And I think, given that inference time compute is so much cheaper than pre-training, the first to achieve AGI might have enough compute on hand from having been to first to build it that they would not need to purchase many more GPU's from Nvidia. So at some point, Nvidia's revenues are going to decline precipitously.
So the question is: how far away are we from AGI? Seems like most experts estimate 3-10 years. Did that timeline just shrink by 50x (or at least by some multiple) from these new optimizations from DeepSeek?
I think Nvidia's revenue is going to continue to grow at its current pace (or faster) until AGI is achieved. Just sell your positions right before that happens.
(Not an expert. Not even an amateur)
That’s spitting distance from “we don’t know”.
I think most of these valuations are made by people without the expertise to predict what the next one year of AI looks like, never mind 3-10 years.
The large investments were mainly for training larger foundation models, or at the very least hedging for that. It hasn't been that clear over the last 1+ years that simply increasing the number of parameters continues to lead to the same improvements we've seen before.
Markets do not necessarily have any prediction power here. People were spooked by DeepSeek getting ahead of the competition and by the costs they report. There is still a lot of work and some of it may still require brute force and more resources (this seems to be true for training the foundation models still).
That they get to be the trillionaires with an untouchable moat? Wouldn't this be like creating a Kwisitz Hadarach thinking you can control it, to borrow a Dune reference?
Maybe this was the point of creating this model all along ?
* Intel, AMD, etc. could start making competitive GPUs that push the price down
* A new ASIC chip specifically designed for LLMs
* A new training or LLM runtime algorithm that uses the CPU
* Quantum chip that can train or run a LLM
If Nvidia lost its AI dominance, where would its stock be?
Around the AMD (200bn) / Intel levels (100bn) which is a ~90% reduction in share price from todays close.
2) to address Nvidia valuation: there is no cap to demand on intelligence (or, we're not close). People will never be satisfied with the intelligence achieved and just stop asking for more. So Nvidia will still sell the hardware as the demand side is uncapped.
Unrelated note that I was considering and would like an opinion on: Nvidia is the software play in AI, and TSMC the hardware play. Nvidia has competitors like broadcom/AMD/TPUs but beats out on software. TSMC will be frontier on manufacturing everyone's hardware.
For Nvidia, it is great news because now finally, the concentration of GPUs at Hyperscalers will end and every Fortune company can finally get their local data center to train their AI Models.
Because if training AI models becomes more efficient and easier then the ones being that business are at risk so basically Big Tech. Nvidia isn't in that business but in the business of providing tools to train.
Fortunately, Big Tech can easily do something to prevent ANYONE for competing. They simply buy all available GPUs. Oh wait, haven't they been doing it for years? Excatly!
People really don't get what an arms race and market competition race is.
How do I prevent disruption? I simply buy all the tools the competition needs to disrupt me.
See, if Fortune 500 companies want to build large data centers but can't because all Hyperscalers buy the GPUs then eventually they will rent from cloud as otherwise they can't get the GPUs.
Spending $50 billion to do $6 million worth of Ai training seems like a good way to trigger a golden parachute and "spend more time with your family" as a CEO.
Everyone is hung up on the cost of training R1 but it’s 685 billion parameters. We still need all of those GPUs to actually use the model.
That makes both NVDA stock and big AI infrastructure spending less compelling, as those needs are scaled down via software efficiency and chip alternatives.
For some time it's been clear that there's an AI bubble and this may be what finally pops it.
The dot com bubble was a real thing, and yet the internet has gone on to be one of humanities most valuable innovations.
* Price has dropped to the level it was at four months ago in October
* The 1-year is up 142%
* The 2-year is up 530%
If you're just going by the market's assessment, this is not a "crash" and the "game has not changed".
Say DeepSeek has worked out how to do more with less - that's great! I don't think it means that the market for Nvidia's silicon (or anyone else who can hop the CUDA moat) is going to shrink. I think that the techniques for doing more with less.. will be applied to _more_ hardware, to make even bigger things. AI is in its infancy, and frankly has a long way to go. An efficiency multiplier at this stage of growth isn't going to reduce the capital needed, it will just make it go further. There may be some concern about scaling the amount of training material, but I don't see that as the end of the road at all. After all, a human's mental growth is hardly limited by the amount of available reading material. After all written material is trained upon, the next frontier will just be in some other mimicry of biological metacognition.
At least that's my armchair analysis.
do models with DeepSeek architecture still scale up?
If yes, then bigger clusters will outperform in near future. NVidia wins as tide rises all boats, and them first.
If not, then it's still possible to run several models in parallel to do the same, potentially big, job. Just like humans team. All we need is to learn how to do it efficiently. This way bigger clusters win again.
I can't seem to find any evience beyond the team's statements. What have I missed?
Compute resources they officially should not access to given export bans, where mentioning them might lead to their export ban bypass getting rolled up.
Hypothetical: Take a large short position on NVDA, announce the market that you trained a massive model without using 10s of millions of rare-as-hen's-teeth NVDA-sentsitive resources. Settle postition then quietly settle giant compute bill. Difficult to know either way, but the market seems have taken the team at face value. I guess we'll know if and only if this reduced training cost methodology is replicated.
The impact of competition and DeepSeek on Nvidia - https://news.ycombinator.com/item?id=42822162 - Jan 2025 (423 comments)
1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Paradox.
2. To build a moat, OpenAI and American AI companies need to up their datacenter spending even more.
3. DeepSeek's breakthrough is in distilling models. You still need a ton of compute to train the foundational model to distill.
4. DeepSeek's conclusion in their paper says more compute is needed for next break through.
5. DeepSeek's model is trained on GPT4o/Sonnet outputs. Again, this reaffirms the fact that in order to take the next step, you need to continue to train better models. Better models will generate better data for next-gen models.
I think DeepSeek hurts OpenAI/Anthropic/Google/Microsoft. I think DeepSeek helps TSMC/Nvidia.
Just try out the standard deepseek-r1 or even the deepseek-r1:1.5B through ollama.
No need for expensive hardware anymore locally. My PC ( without Nvidia card/expensive hardware) runs the a deepseek 1.5 b query fast enough - 2 - 9 seconds until it's finished.
Further more, reasoning models require more tokens. The faster the GPU, the more thinking it can do in a set amount of time. This means the faster the hardware, the smarter the model output. Again, that reinforces the need for faster hardware.
More thinking = smarter models
Faster hardware = more thinking
Therefore, faster hardware = smarter models
A Chinese company coming up with a cheaper alternative to a cutting-edge technology out of nowhere, is an outcome that is hard to predict.
In hindsight, betting on Nvidia maintaining its monopoly on a resource crucial for such an important technology as AI, might not be the best of ideas, but then again, who knows.
meta uses amd mi300x for all inference and which bought 173.000 gpu ( 40% of their total gpu count ) according to that report https://www.cnbc.com/2023/12/06/meta-and-microsoft-to-buy-am...
What you're really trying to purchase is a machine that creates more money than it uses. You need to guess at if that machine will do its job at an arbitrary point in the future, and how well it will do it. Those factors are only loosely correlated with current PE
crickets
That demand is going to dry up any day now, mark my words. Tomorrow, even!
It was worth $50 in Jan 24.
While the drop may sound huge, we're pretty far from a crash here.
Basically, if you look at only the stories, you get a dire picture -- like, "how will they make payroll????"
However, if you look at the actual share price in context, you see . . . a not very interesting event.
Look here:
https://www.macrotrends.net/stocks/charts/NVDA/nvidia/pe-rat...
The graph definitely shows a share price adjustment, but it by no means erases the run-up Nvidia has enjoyed in the last 2 years. All that happened was the stock dropped back to it price of about the middle of last summer.
This may be old news to most of you, but: in markets we also track a figure called the P/E ratio, which is the ratio of the price of the company vs. its profits. Old-school manufacturing firms -- what used to be "blue chip" stocks -- would be in the 18-22 range here. Apple which absolutely PRINTS cash, is at a high-flying 38. NVidia's P/E is still 58.8.
The tl;dr here is that NVidia is still valued very, very highly. They're still what, the 3rd most valuable company in the world, with an enviable position in chipmaking. They still make money hand over fist. The weird part is that their firm is SO valuable that they could take a $600B haircut and it be almost a non-event.
If that's true, normally, mainland will win, as the people over that side have more grit, are more eager to succeed, are working at 996 schedule, and have nothing to lose. They're hard to stop as long as their government will not interfere.
Guess we shall see how it pans out…
https://m.youtube.com/watch?v=7LNyUbii0zw
He seems adamant that there are no diminishing returns to scaling AI.
I don’t want to stir up conspiracy theories but I do think that currently all the big AI players have a vested interest in the message that the current scaling paradigm is the right one, and that this is a supremacy issue wrt China. It drives so much investment and valuation that I doubt they can truly be objective.
500 Billion is a lot of money. Expect even crimes to be commited in order to make it happen.
Why buy 100k Hoppers if 20k Blackwell offer the same compute so then it's better to buy 100k Blackwells right?
Backwell will increase cluster scaling easily by 10x performance and if you buy 10x of them then your compute on a cluster will be 100x than before. If it takes you to wait 6-12 months for that then so be it. You will easily make up the time in the end with the speed up.
Everyone else who over-leveraged into the private AI companies (Anthropic, OpenAI) are going to have their valuations under scrutiny.
It actually affects the frontier AI companies (OpenAI, Anthropic, etc) who directly make money from their closed models AND spend hundreds of millions on training these models.
Why pay $3/per million tokens (Claude 3.5 Sonnet) when DeepSeek R1 offers $0.14 / per million tokens and the model is on par with (OpenAI o1) and R1 itself is released for free?
$0 free AI models are eating closed AI models lunch.
But yeah agree with you!
cheaper spam and impersonation engines (LLMs) have a direct impact on the ad network that is strategically downsizing its moderation efforts
And yet again a cheaper Chinese product turns up and everyone loses their minds. Expect a ban incoming to preserve the AI valuations.
Meanwhile in another tech industry a startup had to think lean and innovate its way around resource restrictions. And OpenAI looks like a mess now.
The RL techniques present will only work in domains where you can guarantee an answer is right (multiple choice questions, math, etc.). It doesn't really present any convincing leap forward in terms of advancing the capability of LLMs, just a strategy for compute efficient distillation of what we know already works. The fact this shitty PPO proxy works at all is a testament to the fact that DeepSeek is bootstrapping its capability heavily off of the output of existing larger models which are much more expensive to train. What DeepSeek R1 proves is you can distill a ChatGPT et al. into a smaller model and hack certain benchmarks with RL.
If you could just do RL to predict the best next word in general this would have been done already - but the signal to noise ratio on exploration would be so bad you'd never get anything besides infinite monkeys at a typewriter. It's not a novel/complicated idea to anyone familiar with RL to try and improve probability of things you like, and whoever decided to do RLHF on an LLM surely thought of (and did) regular RL first - and found it didn't work very well with whatever pretrained model and rewards they had. it was like two weeks ago people were going crazy about O3 doing arc-agi by running the exact same kind of traces R1 is doing in "GRPO" at test time rather than train time. Doing this also isn't novel and also only helps on shitty toy problems where you can get a number to tell you good vs bad.
There is no mechanism to compute rewards for general purpose language tasks - and in fact I think people will come to see the gains in math/coding benchmark problems come at a real cost to other capabilities of models which are harder to quantify and impossible to generically assign rewards to at internet scale.
To explore the frontier of capability you will still need a massive amount of compute, in fact even more to do RL than you would need to do standard next token prediction - even if the LLM might have fewer paramters. You also can't afford to do all the optimizations as you try many different complex architectures.
Making a pure-research, foundation model company is silly. Make a product company that sells products.
Startups don't have that option.
1) Their initial AI offerings weren't real products customers would use or pay for
2) They weren't seeing sufficient adoption to justify the expense
3) They have insane levels of distribution in their existing product lines and can incrementally add AI features
This is entirely orthogonal to whether or not other startups can build AI-first products or whether they can position themselves to compete with the giants.
What moat does an AI startup have that will prevent them from being crushed by big tech?
That's gonna be a looong nap
A more efficient model is better for NVIDIA not worse. More compute is still better given the same model. And as more efficient models proliferate it means more edge computing which means more customers with lower negotiating power than Meta and Google…
This is like thinking that if people need to dig only 1 day instead of an entire month to get to their nugget of gold in the midst of a gold rush the trading post will somehow sell fewer shovels…
And whether it’s overvalued or not isn’t relevant that selling a stock because the product the company produces is now even more effective is mind bogglingly stupid.
If it’s cheaper to train models it means far more customers that will try their luck.
If you reduce training requirement from a 100,000 GPUs to a 1000 you’ve now opened the market to 1000’s and 1000’s of potential players instead of like the 10 that can afford dumping so much money into a compute cluster.
The goal for AGI and ASI MUST BE to train, inference, train, inference and so on and that all on the fly in fractions of a second from every token produced.
Now good luck calculating the compute and hard work in algorithms to get there.
Not possible? Then AGI won't ever work because how can AGI beat a human if it can't learn on the fly? Not to mention ASI lol.
Just for those that clearly have no idea https://old.reddit.com/r/AMD_Stock/comments/1d2okn1/when_wil...
You shouldn't underestimate the fact that a large amount of these trades are on margin. Sometimes you can't wait it out because you'll get margin called and if you can't pony up additional cash you're basically getting caught with your pants down.
Disclaimer: I am not a trader, so could be way off
If it’s cheaper to inference you end up using the model for more task, it it’s cheaper to train you train more than models. And if you now need only 1000’s of GPUs instead of 10’s or 100’s of thousands you’ve just unlocked a massive client base of those who can afford to invest high six to low seven figures instead of 100’s of millions or billions into to try their luck.
They have a lot of very smart people and the will to do it, seems like a matter of time before they succeed.
The proof is in the pudding, you're welcome to prove "everyone" wrong.
No it isn't. Investors are most likely expecting there will be less demand for Nvidia's product long-term due to these alleged increased training efficiencies.
You seem to believe that the more inference or training value per piece of tech the more demand there will be for that piece of tech full stop when there are multiple forces at play. As a simple example, you can think of this as a supply spike; while you can make the bet that the demand will follow there could be a lag on that demand spike due to the time it takes to find use cases with product/market fit. That could collapse prices over the near term which could in turn decrease revenue. As a reminder the stock value isn't a bet on whether "the gold trader" will sell more gold or not, it's a bet on whether the net future returns of the gold trader will occur in line with expectations, expectations that are sky high and have zero competition built in.
So they're in a plushy seat, until the US decides they aren't.
In addition, IMO NVDA’s margins are a gift and a curse. They look great to investors, but also mean all their customers are aggressively looking to produce their own GPUs.
I’m surprised they haven’t yet.
Also, I should add that Deepseek just showed the top GPUs are not necessary to deliver big value.
[1] https://engineering.fb.com/2024/03/12/data-center-engineerin...
This announcement is one step in our ambitious infrastructure roadmap. By the end of 2024, we’re aiming to continue to grow our infrastructure build-out that will include 350,000 NVIDIA H100 GPUs as part of a portfolio that will feature compute power equivalent to nearly 600,000 H100s.
Every single time...
https://en.wikipedia.org/wiki/Jevons_paradox
In economics, the Jevons paradox occurs when technological progress increases the efficiency with which a resource is used, but the falling cost of use induces increases in demand enough that resource use is increased, rather than reduced.
It may be great news for VRAM manufacturers tough.
In other words, NVIDIA is in the red not because the company is suddenly doing worse, but because traders think other traders think it will trade down. That is a self fulfilling prophecy, but only so long as there is sufficient attention to drive that. The same works the other way around as well, so long as there is sufficient attention to drive the AI hype train upwards, related stocks will do well as well.
Well put. People need to unterstand that some stocks are basically one giant casino poker table. There was a comment with a link here that a lot of Nvidia buyers don't even know what products Nvidia is making and they don't care, they just want to buy low and sell high. Insert old famous comment abut shoe shine boy giving investment advice to Wall Street stock traders.
Yesterdays price of (say) NVidia was based on the expectation that companies would need to buy N billion of USD of GPUs per year. Now Deepseek comes out and makes a point that N/10 would be enough. From there it can go two ways:
- NVidia's expected future sales drop by 90%.
- The reduced price for LLMs should allow companies to push AI into markets that were previously not cost effective. Maybe this can 10x the total available market, but since the estimated total available market was already ~everything (due to hype) that seems unlikely.
- NVidia finds another usecase for GPUs to offset the reduced demand from AI companies.
In practice, it will probably be some combination of all three. The real problems are not caused for the "shovel sellers" but for companies like OpenAI and Anthropic, who now suddenly have to compete against a competitor that can produce the same product at (apparently) a fraction of the price.
So if the stock market was reflective of the economy (future or the present) then stocks should go up, instead they're going down. Why? Because the stock market is not reflective of the economy.
The stock market is essentially a reflection of societal perception. DJT which was brought up earlier is a great example, because the price of DJT has next to nothing to do with Trump's businesses and almost everything to do with how he is perceived (and remember there is no such thing as bad publicity).
Personally I think the fall will be momentary and followed shortly by a climb to recovery and beyond, but who really knows.
If you don't want to lose your money: Don't let the sensationalist financial journalists and pundits get to you, don't let big red numbers in your portfolio scare you, ignore traders (they all lose their money), don't sell your stocks unless you actually need that money for something right now, re-read your investment manifesto if you have one, and maybe buy the dip for shits and giggles if you have some spare cash laying around.
OpenAI and Anthropic can react by adopting DeepSeek's compute enhancements and using them to build even better models. AI training is still very clearly compute-limited from their POV (they have more data than they know what to do with already, and training "reasoning"/chains-of-thought requires a lot of reinforcement learning which is especially hard) so any improvement in compute efficiency is great news no matter where it comes from.
Maybe part of the growth was also "stupidness", and in that case buying the dip is a mistake because the "merit" price (value) is still way below.
In the .com bust you could have "bought the dip" in the early 00s right after the crash started and still taken 5 years before you weren't in the red even on "good" (in hindsight) stocks like amazon, ebay, microsoft, etc. The big hype there was eCommerce - it turned out to be true! We use eCommerce all the time now, but it took longer than predicted during the .com boom (same for broadband internet enabling "rich web experience" - it came true, but not fast enough for some hyped companies in '00).
And if you bought some of the darling stocks back then like Yahoo or Netscape that ended up not so great in hindsight you may have never recouped your losses.
There was never a question of if NVDA hardware would have high demand in 2025 and 2026. Everyone still expects them to sell everything they make. The reason the stock is crashing is because Wall St believed that companies who bought 50B+ of NVDA hardware would have a moat. That was obviously always incorrect, TPUs and other hardware was eventually going to be good enough for real world use cases. But Wall St is run by people who don't understand technology.
If they'll sell everything they make and it's all about the moat of their clients, why is NVDA still down 15% premarket? You could quote correlation effects and momentum spillover, but that is still just the higher order effects I mentioned about people's expectations being compounded and thus reactions to adverse news being convex.
Presumably because backorders will go down, production volume and revenue won't grow as fast, Nvidia will be forced to decrease their margins due to lower demand etc. etc.
Selling everything you make is an extremely low bar relative to Nvidia's current valuation because it assumes that Nvidia will be able to grow at a very fast pace AND maintain obscene margins for the next e.g. ~5 years AND will face very limited competition.
So I still don't understand what it is that you are so strongly disagreeing with, and I also don't understand how having owned NVidia stock somehow lends credence to your argument.
We are in agreement that this won't threaten NVidia's immediate bottom line, they'll still sell everything they build, because demand will likely rise to the supply cap even with lower compute requirements. There are probably a multitude of reasons why the very large number of people who own NVidia stock have decided to de-lever on the news, and a lot of it is simple uneducated herding.
But we are fundamentally dealing with a power law here - the forward value expectations for NVidia have exponential growth baked in to the hilt, combined with some good old fashioned tulip mania, and when that exponential growth becomes just slightly less exponential, that results in fairly significant price oscillations today - even though the basic value proposition is still there. This was the gist of my comment - you disagree with this?
Now is looks like that 10x of flow of money into OpenAI will no longer exist. There will be competition and compodiditzation, which causes the value of the tokens to drop way more than 40x.
There has always been a component of gambling to all investing, but that component now seems to utterly eclipse everything else. Merit doesn’t even register. Fundamentals don’t register.
But that's also dumb, because "huge leap forward in training efficiency" is not exactly bad news for the major players in even the medium term. Short term, it means their models are less competitive, but I don't see any reason that they can't leverage e.g. these new mixed precision training techniques on their giant GPU farms and train something even bigger and smarter.
There seems to be this weird baked in assumption that AI is at a permanent (or at least semi-permanent) plateau, and that open source models catching up is the end of the game. But this is an arms race, and we're nowhere near the finish line.
But I agree in the sense that Deepseek just creates more demand. Because people desire to get AI to do more work. This makes bang for buck greater opening new opportunities.
This sell off is like selling Intel in 2010 because of a new C compiler.
unless it can be said we need more performance than is currently possible, e.g. new demand, it would be catastrophic. it is unclear that throwing more compute actually expands what is possible. if that is not the case, efficiency is bad for nvidia because it simply results in less demand.
Today a CPU setup is still nowhere near as fast as a GPU setup (for ML/AI), but who knows how it looks like in the future.
> it is unclear that throwing more compute actually expands what is possible
Wasn't that demonstrated to be true already in the GPT1/2 days? AFAIK, LLMs became a thing very much because OpenAI "discovered" that "throwing more compute (and training data) at the problem/solution expands what is possible"
90% of traders lose money, so that's a data point...
You're trying to apply rational thinking but that's not how markets work. In the end valuations are more about narratives in the collective mind than technological merit.
You answered your own question. People do not dig in the Sacramento right anymore for gold, because, it is gone. If you can train models for 1/100 the cost, and you sell model training chips, you probably are not going to sell as many chips.
Everyone here thinks Nvidia is dommed because of training efficiency.
But what has Nvidia been doing for the past decade? Correct increasing training and inferencing efficiency by magnitudes.
Try to train GPT4 on 10k of Volta, Ampere, Hopper and then Blackwell.
What has happened since then? Nvidia has increased their sales in magnitudes.
Why? Because thanks to improvement in data, in algorithms, compute efficiency ChatGPT was possible in the first place.
Imagine Nvidia wouldn't exist. When do you think the ChatGPT moment would happen on CPUs? LOL
Going back to my first sentence. Nvidia started also with small shovels which were GeForce cards with CUDA. Today Nvidia is selling huge GPU clusters (mining machines, yes pun intended ^^).
No. The stock is still x10 after this dip from 2 years ago and x40 from a few years ago.
Likely a "how solid is the technical moat" evaluation - this could be a one-off or could be that there are an avalanche of advancements to continue along the efficiency side of the process.
Given the style and hype of logic in the AI space, I fully believe resources are not well allocated in compute and _actual_ thinking as to how they are spent.
Deepseek's apparent 10x more efficient per inference token... implies a lot of other hardware meets the general use-case. We also know that reasoning should be about 10W for human speed-of-thought... maybe another 1-2 orders of power efficiency.
"Pre-Training: Towards Ultimate Training Efficiency
We design an FP8 mixed precision training framework and, for the first time, validate the feasibility and effectiveness of FP8 training on an extremely large-scale model. Through co-design of algorithms, frameworks, and hardware, we overcome the communication bottleneck in cross-node MoE training, nearly achieving full computation-communication overlap. This significantly enhances our training efficiency and reduces the training costs, enabling us to further scale up the model size without additional overhead. At an economical cost of only 2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model. The subsequent training stages after pre-training require only 0.1M GPU hours." [1]
I believe the saying is "The market can stay wrong for longer than you can stay solvent."
There is a point where there are enough shovels circulating that the demand for new shovels falters, even with zero drawback in the rush. And if so much gold was being mined that it overwhelmed the market and reduced the commodity price, the value of better shovels is reduced.
DeepSeek and friends basically reduce the commodity value of AI (and to be fair, Facebook, Microsoft et al are trying to do the same thing with their open source models, trying to chop the legs out of the upstart AI cos). If AI is worth less, there are going to be fewer mega capitalized AI ventures buying a trillion dollars worth of rapidly-depreciating GPUs in hopes at eeking out some minor advantage.
I wouldn't short nvidia stock, but at the same time there is a point where the spend of GPUs just isn't rational anymore.
>And as more efficient models proliferate it means more edge computing which means more customers with lower negotiating power than Meta and Google
Edge compute has infinitely more competition than the data center.
Anyway I'm homeless so the not eating all day maybe kinda possibly having something to do with wiping out 600 billion dollars in stock market valuation just on the off chance that it might based on nothing more than wishful thinking?
Nope. Not eating tomorrow. Wonder what will happen.
And then - something completely unexpected, a total curveball - arrives a week into office and everything changes. Your agenda collapses and you enter reactive mode.
AI bubble could pop. Valuations drop. Crypto might get hit. Where is Project Stargate now? It might become a joke.
A more cynical observer might suggest that _the timing was no coincidence_.
Where has Trump's agenda collapsed? I might have missed the press release.
And why would a curveball on AI throw off an agenda around trade, immigration and military engagement? I don't follow.
China could take the lead on AI and I don't see it would impact any of those things. Isn't DeepSeek open source? The US already has access to it, so what leverage could China possible have?
> A more cynical observer might suggest that _the timing was no coincidence_.
Do you feel the Chinese government closely controls AI research and timed this response?
Any evidence for that?
Well, it's too early to say if it has happened in this case. But we've seen it happen again and again, so it will not be a surprise if it happens.
> And why would a curveball on AI throw off an agenda around trade, immigration and military engagement? I don't follow.
Well, trade restrictions against China have just backfired spectacularly. So further trade restrictions may not seem as good an idea. And to start trade wars, you need a strong economy. And at the moment America's economy is entirely driven by the AI bubble, as it is the value of the AI stocks that has separated the economy from the European trend.
It is very likely that military engagement will be driven by, or affected by, developments in AI. And there's no doubt that Taiwan's situation is heavily affected by chip production.
> Do you feel the Chinese government closely controls AI research and timed this response?
No evidence, but I am not naive enough to think it's out of the range of possibility. I don't follow your reasoning, more likely the timing of the release is all they had to modify - which is quite trivial for any government. To not question that would be very naive. They have _every motive_.
Ah ok, so you’re just guessing. You wrote it as if it had already happened.
And I think you’re confusing the stock market with the broader economy. The economy as a whole is unaffected by AI at this point.
And the trade war with China involved hundreds of industries beyond AI. Most of it is manufacturing. I’m not sure how failure of AI sanctions (questionable conclusion) somehow means the trade issues around machine parts needs to be abandoned.
I agree China has motive but I’ve heard so many claims that the Chinese government doesn’t control research or businesses like TikTok.
That's not a generous reading. I'm saying this is what has happened historically.
Edit: the use of the words "might" and "could" made this pretty clear.
> The economy as a whole is unaffected by AI at this point.
True in terms of no marked effect on GDP. But the stock market and the broader economy are very much linked. See 2008. And the stock market is high on AI.
> I’ve heard so many claims that the Chinese government doesn’t control research or businesses like TikTok.
Even western countries tell tech companies what to do. Do you think an authoritarian government is going to do that more or less? There's a reason they didn't want to sell TikTok.
Deepseek showing that you can do pure online RL for LLMs means we now have a clear path to just keep throwing more compute at the problem! If anything we made the whole "we are hitting a data wall" problem even smaller.
Additionally, its yet another proof point that scaling inference compute is a way forward. Models that think for hours or days are the future.
As we move further into the regime of long sequence inference, compute scales by the square of the sequence length.
The lesson here was not "training is going to be cheaper than we thought". It's "we must construct additional pylons uhhh _PUs"
Markets remain irrational and all that...
This market doesn't make any sense.
The SFT training data is hard to produce, while the RL they used was fairly uncomplicated heuristic evaluations and not a secondary critic model. So their RL is a simple approach.
If I’ve said anything wrong, feel free to correct me.
Ultimately someone in America will get desperate enough and start a war when they still have a chance to win. See the Earth-Mars conflict in the Expanse.