The impact of competition and DeepSeek on Nvidia
youtubetranscriptoptimizer.com
youtubetranscriptoptimizer.com
Nvidia’s $589B DeepSeek rout - https://news.ycombinator.com/item?id=42839650 - Jan 2025 (574 comments)
Back then a really big motivator for Asynchronous Transfer Mode (ATM) and fiber-to-the-home was the promise of video on demand, which was a huge market in comparison to the Internet of the day. Just about all the work in this area ignored the potential of advanced video coding algorithms, and assumed that broadcast TV-quality video would require about 50x more bandwidth than today's SD Netflix videos, and 6x more than 4K.
What made video on the Internet possible wasn't a faster Internet, although the 10-20x increase every decade certainly helped - it was smarter algorithms that used orders of magnitude less bandwidth. In the case of AI, GPUs keep getting faster, but it's going to take a hell of a long time to achieve a 10x improvement in performance per cm^2 of silicon. Vastly improved training/inference algorithms may or may not be possible (DeepSeek seems to indicate the answer is "may") but there's no physical limit preventing them from being discovered, and the disruption when someone invents a new algorithm can be nearly immediate.
The rise of the net is Jevons paradox fulfilled. The orders of magnitude less bandwidth needed per cat video drove much more than that in overall growth in demand for said videos. During the dotcom bubble's collapse, bandwidth use kept going up.
Even if there is a near-term bear case for NVDA (dotcom bubble/bust), history indicates a bull case for the sector overall and related investments such as utilities (the entire history of the tech sector from 1995 to today).
Plus the cell phone industry paved the way for VOIP by getting everyone used to really, really crappy voice quality. Generations of Bell Labs and Bellcore engineers would rather have resigned than be subjected to what's considered acceptable voice quality nowadays...
HLS is really just a way to empower the client with the ownership of the playback logic. Let the client handle forward buffering, retries, stream selection, etc.
What accounts for this difference? Is there something inherently worse about the nature of cell phone infrastructure over land-line use?
I'm totally naive on such subjects.
I'm just old enough to remember landlines being widespread, but nearly all of my phone calls have been via cell since the mid 00s, so I can't judge quality differences given the time that's passed.
Has it not been like this for a very long time? I was under the impression that "voice frequency" being defined as up to 4 kHz was a very old standard - after all, (long-distance) phone calls have always been multiplexed through coaxial or microwave links. And it follows that 8kbps is all you need to losslessly digitally sample that.
I assumed it was jitter and such that lead to lower quality of VoIP/cellular, but that's a total guess. Along with maybe compression algorithms that try to squeeze the stream even tighter than 8kbps? But I wouldn't have figured it was the 8kHz sample rate at fault, right?
Huh? What? That's not even remotely true.
If you read your comment out loud, the very first sound you'd make would have almost all of its energy concentrated between 4 and 10 kHz.
Human vocal cords constantly hit up to around 10 kHz, though auditory distinctiveness is more concentrated below 4 kHz. It is unevenly distributed though, with sounds like <s> and <sh> being (infamously) severely degraded by a 4 kHz cut-off.
The speech codecs are complex and fascinating, very different from just doing a frequency filter and compressing.
The base is linear predictive coding, which encodes the voice based on a simple model of the human mouth and throat. Huge compression but it sounds terrible. Then you take the error between the original signal and the LPC encoded signal, this waveform is compressed heavily but more conventionally and transmitted along with the LPC signal.
Phones also layer on voice activity detection, when you aren't talking the system just transmits noise parameters and the other end hears some tailored white noise. As phone calls typically have one person speaking at a time and there are frequent pauses in speech this is a huge win. But it also makes mistakes, especially in noisy environments (like call centers, voice calls are the business, why are they so bad?). When this happens the system becomes unintelligible because it isn't even trying to encode the voice.
The 8kbps rates on cellular are the more complicated (relative to G.711) AMR-NB encoding. AMR supports voice rates from about 5-12kbps with a typical 8kbps rate. There's a lot more pre and post processing of the input signal and more involved encoding. There's a bit more voice information dropped by the encoder.
Part of the quality problem even today with VoLTE is different carriers support different profiles and calls between carriers will often drop down to the lowest common codec which is usually AMR-NB. There's higher bitrate and better codecs available in the standard but they're implemented differently by different carriers for shitty cellular carrier reasons.
I'm a moron, thanks. I think I got the sample rate mixed up with the bitrate. Appreciate you clearing that up - and the other info!
1. it takes considerable work on my part to understand it on a cell phone
2. it's much easier on POTS
3. it's not a problem on VOIP
4. no issues in person
With all the amazing advances in cell phones, the voice quality of cellular is stuck in the 90's.
At home, I use VoLTE and the sound is almost impeccable, very high quality, but in the places I roam to, what I get is FM quality 3G sound.
It's not that cellular network is incapable of that sound quality, but I don't get to experience it except my home country. Interesting, indeed.
3G networks in many European countries were shut off in 2022-2024. The few remaining ones will go too over the next couple of years.
VoLTE is 5G, common throughout Europe. However the handset manufacturer may need to qualify each handset model with local carriers before they will connect using VoLTE. As I understand the situation, Google for instance has only qualified Pixel phones for 5G in 19 of 170-odd countries. So 5G features like VoLTE may not be available in all countries. This is very handset/country/carrier-dependent.
Technically, on 5G you have "VoNR"[0], where VoLTE is over 4G.
Also, the phone companies had a pathological aversion to understanding Moore's law, because it suggested they'd have to charge half as much for bandwidth every 18 months. Long distance rates had gone down more like 50%/decade, and even that was too fast.
Better video compression led to an explosion in video consumption on the Internet, leading to much more revenue for companies like Comcast, Google, T-Mobile, Verizon, etc.
More efficient LLMs lead to much more AI usage. Nvidia, TSMC, etc will benefit.
It’s very shortsighted to think we’re going to need fewer chips because the algorithms got better. The system became more efficient, which causes induced demand.
Most likely, consumer surplus has gone up.
Anyone currently invested is presumably in because they like the insanely high profit margin, and this is apt to quash that. There is now much less reason to give your first born to get your hands on their wares. Comcast, Google, T-Mobile, Verizon, etc., and especially those not named Google, have nothingburger margins in comparison.
If you are interested in what they can do with volume, then there is still a lot of potential. They may even be more profitable on that end than a margin play could ever hope for. But that interest is probably not from the same person who currently owns the stock, it being a change in territory, and there is apt to be a lot of instability as stock changes hands from the one group to the next.
I'm invested in Nvidia because it's part of the index that my ETF is tracking. I have no clue what their profit margins are.
That would be an unusual situation for an ETF. An ETF does not usually extend ownership of the underlying investment portfolio. An ETF normally offers investors the opportunity to invest in the ETF itself. The ETF is what you would be invested in. Your concern as an investor in an ETF would only be with the properties of the ETF, it being what you are invested in, and this seems to be true in your case as well given how you describe it.
Are you certain you are invested in Nvidia? The outcome of the ETF may depend on Nvidia, but it may also depend on how a butterfly in Africa happens to flap its wings. You aren't, by any common definition found within this type of context, invested in that butterfly.
The ETF is just one more layer of indirection. You might like to read https://en.wikipedia.org/wiki/Exchange-traded_fund#Arbitrage... to see how ETFs are connected to the underlying assets.
You will find that the connection between ETFs and the underlying assets in the index is much more like the connection between your Robinhood portfolio and Nvidia, than the connection between butterflies and thunderstorms.
[0] At least for its stocks. Its bonds are probably held in different but equally weird ways.
Technically, but they extend ownership. An ETF is a different type of abstraction. Which you already know because you spoke about that abstraction in your original comment, so why play stupid now?
An ETF typically holds the underlying assets, and you own a part of the ETF.
If the AI market gets 10x bigger, and GPU work gets 50% smaller (which is still 5x larger than today) - but Nvidia is priced on 40% growth for the next ten years (28x larger) - there is a price mismatch.
It is theoretically possible for a massive reduction in GPU usage or shift from GPU to CPU to benefit Nvidia if that causes the market to grow enough - but it seems unlikely.
Also, I believe (someone please correct if wrong) DeepSeek is claiming a 95% overall reduction in GPU usage compared to traditional methods (not the 50% in the example above).
If true, that is a death knell for Nvidia's growth story after the current contracts end.
Nope. Anything inheriantly serial is better off on the CPU due to caching and it's architecture.
Many things that are highly parallizable are getting GPU enabled. Games and ML are GPU by default, but many things are migrating to CUDA.
You need both for cheap, high performance computing. They are different workloads.
No - because this eliminates entirely or shifts the majority of work from GPU to CPU - and Nvidia does not sell CPUs.
I'm not even sure how to reply to this. GPUs are fundamentally much more efficient for AI inference than CPUs.Whether your new coprocessor or instructions look more like a GPU or something else doesn't really matter if we are done squinting and calling it graphics like problems and/or claiming it needs a lot more than a middle class PC.
CPU chipsets have borrowed video decoder units and SSE instructions from GPU-land, but the idea that video decoding is a generic CPU task now is not really true.
Now maybe every computer will come with an integrated NPU and it won't be made by Nvidia, although so far integrated GPUs haven't supplanted discrete ones.
I tend to think today's state-of-the-art models are ... not very bright, so it might be a bit premature to say "640B parameters ought to be enough for anybody" or that people won't pay more for high-end dedicated hardware.
Depends on what form factor you are looking at. The majority of computers these days are smart phones, and they are dominated by systems-on-a-chip.
Not only are 10-100x changes disruptive, but the players who don't adopt them quickly are going to be the ones who continue to buy huge amounts of hardware to pursue old approaches, and it's hard for incumbent vendors to avoid catering to their needs, up until it's too late.
When everyone gets up off the ground after the play is over, Nvidia might still be holding the ball but it might just as easily be someone else.
If stock market-cap is (roughly) the market's aggregated best guess of future profits integrated over all time, discounted back to the present at some (the market's best guess of the future?) rate, then increasing uncertainty about the predicted profits 5-10 years from now can have enormous influence on the stock. Does NVDA have an AWS within it now?
Cisco in 1994: $3.
Cisco after dotcom bubble: $13.
So is Nvidia's stock price closer to 1994 or 2001?
Bandwidth is one thing, but the real benefit is that ATM also guaranteed minimal latencies. You could now shave off another 20-100ms of latency for your FaceTime calls, which is subtle but game changing. Just instant-on high def video communications, as if it were on closed circuits to the next room.
For the same reasons, the AI analogy could benefit from both huge processing as well as stronger algorithms.
Which means you need state (and the overhead that goes with it) for each connection within the network. That's horribly inefficient, and precisely the reason packet-switching won.
> An internet based on ATM would have been amazing.
No, we'd most likely be paying by the socket connection (as somebody has to pay for that state keeping overhead), which sounds horrible.
> You could now shave off another 20-100ms of latency for your FaceTime calls, which is subtle but game changing.
Maybe on congested Wi-Fi (where even circuit switching would struggle) or poorly managed networks (including shitty ISP-supplied routers suffering from horrendous bufferbloat). Definitely not on the majority of networks I've used in the past years.
> The horror of packet switching is all the buffering it needs [...]
The ideal buffer size is exactly the bandwidth-delay product. That's really not a concern these days anymore. If anything, buffers are much too large, causing unnecessary latency; that's where bufferbloat-aware scheduling comes in.
The latency benefit would outweigh the cost. Just absolutely instant video interaction.
And for the little bit of impact queueing latency has (if done well, i.e. no bufferbloat), I doubt anyone would notice the difference, honestly.
What I recall is that it was at a time when Internet folks had made enormous advances in understanding congestion behavior in computer networks, and other folks (e.g. my division of Motorola) had put a lot of time into understanding the limited burstiness you get with silence suppression for packetized voice, and these folks knew nothing about it.
You don't actually need all that much buffering.
Buffer bloat is actually a big problem with conventional TCP. See eg https://news.ycombinator.com/item?id=14298576
I already do this. But I cheat - I use a good router (OpenWrt One) that has built-in controls for Bufferbloat. See [How OpenWrt Vanquishes Bufferbloat](https://forum.openwrt.org/t/how-openwrt-vanquishes-bufferblo...)
DCT was developed in 1972 and has a compression ratio of 100:1.
H.264 compresses 2000:1.
And standard resolution (480p) is ~1/30th the resolution of 4k.
---
I.e. Standard resolution with DCT is smaller than 4k with H.264.
Even high-definition (720p) with DCT is only twice the bandwidth of 4k H.264.
Modern compression has allowed us to add a bunch more pixels, but it was hardly a requirement for internet video.
Modern compression algorithms were developed but not even computationally available for some of the time.
It doesn’t have a compression ratio.
Therefore, it has no compression ratio, and it doesn’t make sense to compare it to other algorithms.
The real compression comes from quantization and entropy coding (Huffman coding, arithmetic coding, etc.).
Didn't TMSC say that SamA came for a visit and said they needed $7T in investment to keep up with the pending demand needs.
This stuff is all super cool and fun to play with, I'm not a nay sayer but it almost feels like these current models are "bubble sort" and who knows how it will look if "quicksort" for them becomes invented.
That only lasted so long. Then it was heavy machinery (hydraulics, excavators, etc)
As pointed out in the article, Nvidia has several advantages including:
- Better Linux drivers than AMD
- CUDA
- pytorch is optimized for Nvidia
- High-speed interconnect
Each of the advantages is under attack: - George Hotz is making better drivers for AMD
- MLX, Triton, JAX: Higher level abstractions that compile down to CUDA
- Cerbras and Groq solve the interconnect problem
The article concludes that NVIDIA faces an unprecedented convergence of competitive threats. The flaw in the analysis is that these threats are not unified. Any serious competitor must address ALL of Nvidia's advantages. Instead Nvidia is being attacked by multiple disconnected competitors, and each of those competitors is only attacking one Nvidia advantage at a time. Even if each of those attacks are individually successful, Nvidia will remain the only company that has ALL of the advantages.* Groq can't produce more hardware past their "demo". It seems like they haven't grown capacity in the years since they announced, and they switched to a complete SaaS model and don't even sell hardware anymore.
* I dont know enough about MLX, Triton, and JAX,
I think you should consider it as, if they're trying to avoid Nvidia and make sure their code isn't tied to NVidia-isms, and AMD is troublesome enough for basics the step to customized solutions is small enough to be worthwhile for something even cheaper than AMD.
Disclaimer: long AMD, and not precise on percentages. Just illustrating a point.
Funny timing though, today NVDA lost $589 billion in market cap as the market got spooked.
ah ... no ... that's nonsense trying to hide behind stilted math lingo.
This does not match my experience from the past ~6 years of using AMD graphics on Linux. Maybe things are different with AI/Compute, I've never messed with that, but in terms of normal consumer stuff the experience of using AMD is vastly superior than trying to deal with Nvidia's out-of-tree drivers.
lol
What came out of it (and the semianalysis article) was that Anush would step up to the plate and work on improving the software.
George making noise is just a momentary blip in time that will be forgotten a week later…
The same is true of electricity - neither nuclear power nor fusion will not be online anytime soon.
Not nearly all data centers are water cooled, and there is this amazing technology that can convert sunlight into electricity in a relatively straightforward way.
AI workloads (at least training) are just about as geographically distributeable as it gets due to not being very latency-sensitive, and even if you can't obtain sufficient grid interconnection or buffer storage, you can always leave them idle at night.
Stock price is based on future earnings.
The smart money knows this and is reacting this morning - thus the drop in NVDA.
Current world marketed energy consumption is about 18 terawatts. Current mainstream solar panels are 21% efficient. At this efficiency, the terrestrial solar resource is about 37000 terawatts, 2000 times larger than the entire human economy:
~ $ units
Currency exchange rates from exchangerate-api.com (USD base) on 2024-11-25
Consumer price index data from US BLS, 2024-11-24
7290 units, 125 prefixes, 169 nonlinear units
You have: 21% solarirradiance circlearea(earthradius)
You want: TW
* 36531.475
/ 2.7373655e-05
IEA reports that currently (three years ago) datacenters used 460TWh/year. In SI units, that's 0.05 terawatts. https://iea.blob.core.windows.net/assets/6b2fd954-2017-408e-...So, once datacenters are using seven hundred thousand times more power than currently, we might need to seek power sources for them other than terrestrial solar panels running microgrids. Solar panels in space, for example.
You could be forgiven for wondering why this enormous resource has taken so long to tap into and why the power grid is still largely fossil-fuel-powered. The answer is that building fossil fuel plants only costs on the order of US$1–4 per watt (either nameplate or average), and until the last few years, solar panels cost so much more than that that even free "fuel" wasn't enough to make them economically competitive. See https://www.eia.gov/analysis/studies/powerplants/capitalcost... for example.
Today, however, solar panels cost US$0.10 per peak watt, which works out to about US$0.35 to US$1 per average watt, depending largely on latitude. This is 25% lower than the price of even a year ago and a third of the price of two years ago. https://www.solarserver.de/photovoltaik-preis-pv-modul-preis...
He says this and talks about it in The Fallout section - even at BigCos with megabucks the teams are starved for time on the Nvidia chips and if these innovations work other teams will use them and then boom Nvidia's moat is truncated somehow which doesn't look good at such lofty multiples
(Famous for hacking the PS3–except he just took credit for a separate group’s work. And for making a self-driving car in his garage—except oh wait that didn’t happen either.)
Why do you think otherwise? Can you share specific details?
Not really, his article focuses on Nvidia's being valued so highly by stock markets, he's not saying that Nvidia's destined to lose its advantage in the space in the short term.
In any case, I also think that the likes of MSFT/AMZN/etc will be able to reduce their capex spending eventually by being able to work on a well integrated stack on their own.
Nvidia are doing phenomenal things with robotics, and that is likely to be the next shoe to drop, and they are positioned for another catalytic moment similar to that which we have seen with LLMS.
I do think we will see some drawback or at least deceleration this year while the current situation settles in, but within the next three years I think we will see humanoid robots popping up all over the place, particularly as labour shortages arise due to political trends - and somebody is going to have to provide the compute, both local and cloud, and the vision, movement, and other models. People will turn to the sensible and known choice.
So yeah, what you say is true, but I don’t think is going to have an impact on the trajectory of nvidia.
Unless something radically changed in the last couple years, I am not sure where you got this from? (I am specifically talking about GPUs for computer usage rather than training/inference)
This was the first thing that stuck out to me when I skimmed the article, and the reason I decided to invest the time reading it all. I can tell the author knows his shit and isn't just parroting everyone's praise for AMD Linux drivers.
> (I am specifically talking about GPUs for computer usage rather than training/inference)
Same here. I suffered through the Vega 64 after everyone said how great it is. So many AMD-specific driver bugs, AMD driver devs not wanting to fix them for non-technical reasons, so many hard-locks when using less popular software.
The only complaints about Nvidia drivers I found were "it's proprietary" and "you have to rebuild the modules when you update the kernel" or "doesn't work with wayland".
I'd hesitate to ever touch an AMD GPU again after my experience with it, haven't had a single hick-up for years after switching to Nvidia.
This isn’t a barrier for Linux veterans but it adds significant resistance for part-time users, even those that are technically inclined, compared to the “it just works” experience one gets with an Intel/AMD GPU under just about every Linux distro.
goat.
Cerbras and Groq need to solve the memory problem. They can't scale without adding 10x the hardware.
In which way? As a user who switched from an AMD-GPU to Nvidia-GPU, I can only report a continued amount of problems with NVIDIAs proprietary driver, and none with AMD. Is this maybe about the open source-drivers or usage for AI?
When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions:
1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability)
2) Figure out how to build a real competitive moat through one or more of the following: economies of scale, network effects, regulatory capture
On the second point, it seems to me like the only realistic strategy for companies like OpenAI is to turn themselves into a platform that benefits from direct network effects. Whether that's actually feasible is another question.
you are assuming that what DeepSeek achieved can be reasonably easily replicated by other companies. then the question is when all big techs and tons of startups in China and the US are involved, how come none of those companies succeeded?
deepseek is unique.
It think they think American engineering excellence was due to neoliberal inginuenity visavi the USSR, not the engineers and the transfer of academic legacy from generation to generation.
I see two possibilities here, either that the CCP is not that all-reaching as we think, or that the value of the technology isn't critical, and that the release was further cleared with the CCP and maybe even timed to come right after Trump's announcement of American AI supremacy.
In a way, their strategy could be:
1) Let the US invest $1 trillion in R&D
2) Support the open source community such that their capability to replicate these models only marginally lags the private sector
3) When R&D costs are more manageable, lean in and play catch up
And it did a great job. Nvidia stock's sunk, and investors are going to be asking if it's really that smart to give American AI companies their money when the Chinese can do something similar for significantly less money.
First mover advantage acquired and keeps subscribers.
No one really cares if you matched GPT4o one year later. OpenAI has had a full year to optimize the model, build tools around the model, and used the model to generate better data for their next generation foundational model.
Why do you expect OpenAI to become profitable after 3 years of chatgpt?
And more important (for us), let the hiring frenzy start again :)
In this world, R&D costs and gross margin/revenue are inextricably correlated.
The market always consolidates when it matures. Every time. The market always consolidates into 2-3 big players. Often a duopoly. OpenAI is trying to be one of the two or three companies left standing.
Brands are incredibly powerful when talking about consumer goods.
Your analogy is valid at this time, but proves the GP's point, not yours.
1) There was a data flywheel effect, wherein Google was able to improve search results by analyzing the vast amount of user activity on its site.
2) There were real economies of scale in managing the cost of data centers and servers
3) Their advertising business model benefited from network effects, wherein advertisers don't want to bother giving money to a search engine with a much smaller user base. This profitability funded R&D that competitors couldn't match.
There are probably more that I'm missing, but I think the primary takeaway is that Google's scale, in and of itself, led to a better product.
Can the same be said for OpenAI? I can't think of any strong economies of scale or network effects for them, but maybe I'm missing something. Put another way, how does OpenAI's product or business model get significantly better as more people use their service?
Their SOTA models can generate better synthetic data for the next training run - leading to a flywheel effect?
A similar dynamic occurred in the early days of search engines.
In practice I never heard OpenAI mention how they use chat logs for improving the model. They are either afraid to say, for privacy reasons, or want to keep it secret for technical advantage. But just think about the billions of sessions per month. A large number of them contain extensive problem solving. So the LLMs can collect experience, and use it to improve problem solving. This makes them into a flywheel of human experience.
1) Google copied the hotmail model of strapping commodity PC components to cheap boards and building software to deal with complexity.
2) Yahoo had a much larger cage, filled with very very expensive and large DEC machines, with one poor guy sitting in a desk in there almost full time rebooting the systems etc....I hope he has any hearing left today.
3) Just right before the .com crash, I was in a cage next to Google's racking dozens of brand new Netra T1s, which were pretty slow and expensive...that company I was working for died in the crash.
Look at Google's web page:
https://www.webdesignmuseum.org/gallery/google-1999
Compare that to Yahoo:
https://www.webdesignmuseum.org/gallery/yahoo-in-1999
Or the company they originaly tried to sell google to Excite:
https://www.webdesignmuseum.org/gallery/excite-2001
Google grew to be profitable because they controlled costs, invested in software vs service contracts and enterprise gear, had a simple non-intrusive text based ad model etc...
Most of what you mention above was well after that model focused on users and thrift allowed them to scale and is survivorship bias. Internal incentives that directed capitol expenditures to meet the mission vs protect peoples back was absolutely a related to their survival.
Even though it was a metasearch, my personal preference was SavvySearch until it was bought and killed or what ever that story way.
OpenAI is far more like Yahoo than Google.
I opted for a fanless graphics board, for just that reason.
AdSense
Claude has effectively no eyeballs. API calls != eyeballs.
Late in 2024, OpenAI had $3.7b in revenue. Meanwhile, Claude’s mobile app hit $1 million in revenue around the same time.
Where do they report these ?
edit i found it here https://www.cnbc.com/2024/09/27/openai-sees-5-billion-loss-t...
"OpenAI sees roughly $5 billion loss this year on $3.7 billion in revenue"
This broken record again.
Just observe reality. OpenAI is leading, by far.
All these "OpenAI has no moat" arguments will only make sense whenever there's a material, observable (as in not imaginary), shift on their market share.
ChatGPT is somewhat less censored (certainly on topics painful to the CCP), and GPT is multi-modal, which is a big selling point.
Depends on your use-case, of course.
The same one that underpins the entire existence of a little company called Spotify: I'm just too lazy to cancel my subscription and move to a newer player.
ChatGPT is still vastly more popular than other, similar chat bots.
Does it? As a chat-based (Claude Pro, ChatGPT Plus etc.) user, LLMs have zero stickiness to me right now, and the APIs hardly can be called moats either.
Theres a lot more still to unpack and I don’t expect this to stay solely in the tech realm. Seems to politically sensitive.
It's also worth keeping in mind that depending on benchmark, these values change (and can shrink quite a bit)
And it's also worth keeping in mind that the drastic drop in training cost(if reproducible) will mean that training is suddenly affordable for a much larger number of organizations.
I'm not sure the impact on GPU demand will be as big as people assume.
In the long run, yes, they will be cheaper due to more competition and better tech. But next month? It will be more expensive.
It'll be much harder to convince people to buy the latest and greatest with this out there.
The usage of existing but cheaper nvidia chips to make models of similar quality is the main takeaway.
So why not buy a more expensive Nvidia chip to run a better model?The DeepSeek R1 model people are freaking out about, runs better with more compute because it's a chain of thoughts model.
The analogous costs would be what OpenAI spent to go from GPT 4 to GPT 4o (i.e., to develop the reasoning model from the most up-to-date LLM model). $5 million is still less than what OpenAI spent but it's not a magnitude lower. (OpenAI spent up to $100 million on GPT4 but a fraction of that to get GPT 4o. Will update comment if I can find numbers for 4o before edit window closes)
>Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely tough problems that require "genius level" reasoning to avoid all the potential pitfalls that would lead a regular LLM astray.
I think this is the most interesting part. We always knew a huge fraction of the compute would be on inference rather than training, but it feels like the newest developments is pushing this even further towards inference.
Combine that with the fact that you can run the full R1 (680B) distributed on 3 consumer computers [1].
If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs, what happens when that is only important to a small fraction of the overall AI compute?
Imagine having 300. Could you build even better models? Is DeepSeek the right team to deliver that, or can OpenAI, Meta, HF, etc. adapt?
Going to be an interesting few months on the market. I think OpenAI lost a LOT in the board fiasco. I am bullish on HF. I anticipate Meta will lose folks to brain drain in response to management equivocation around company values. I don't put much stock into Google or Microsoft's AI capabilities, they are the new IBMs and are no longer innovating except at obvious margins.
I don't pretend to know much about the minutiae of LLM training, but it wouldn't surprise me at all if throwing massively more GPUs at this particular training paradigm only produces marginal increases in output quality.
A human googles "how much does a tire cost?"
They pick out a website from search results, then nav within it to the correct product page and maybe scroll until the price is visible on screen.
Google captures a lot of that data on third party sites. From Perplexity:
Google Analytics: If the website uses Google Analytics, Google can collect data about user behavior on that site, including page views, time on site, and user flow.
Google Ads: Websites using Google Ads may allow Google to track user interactions for ad targeting and conversion tracking.
Other Google Services: Sites implementing services like Google Tag Manager or using embedded YouTube videos may provide additional tracking opportunities
So you can imagine that Google has a kajillion training examples that go: search query (which implies task) -> pick webpage -> actions within webpage -> user stops (success), or user backs off site/tries different query (failure)
You can imagine that even if an AI agent is super efficient, it still needs to learn how to formulate queries, pick out a site to visit, nav through the site, do all that same stuff to perform tasks. Google's dataset is perfect for this, huge, and unparalleled.
It seems like there is MUCH to gain by migrating to this approach - and it theoretically should not cost more to switch to that approach than vs the rewards to reap.
I expect all the major players are already working full-steam to incorporate this into their stacks as quickly as possible.
IMO, this seems incredibly bad to Nvidia, and incredibly good to everyone else.
I don't think this seems particularly bad for ChatGPT. They've built a strong brand. This should just help them reduce - by far - one of their largest expenses.
They'll have a slight disadvantage to say Google - who can much more easily switch from GPU to CPU. ChatGPT could have some growing pains there. Google would not.
Often expenses like that are keeping your competitors away.
This is a step function in terms of efficiency (which presumably will be incorporated into ChatGPT within months), but not in terms of end user experience. It's only slightly better there.
Would it not be useful to have multiple independent AIs observing and interacting to build a model of the world? I'm thinking something roughly like the "councelors" in the Civilization games, giving defense/economic/cultural advice, but generalized over any goal-oriented scenario (and including one to take the "user" role). A group of AIs with specific roles interacting with each other seems like a good area to explore, especially now given the downward scalability of LLMs.
Offtopic, but your comment finally pushed me over the edge to semantic satiation [1] regarding the word "moat". It is incredible how this word turned up a short while ago and now it seems to be a key ingredient of every second comment.
I’m sure if I looked, I could find quotes from Warren Buffet (the recognized originator of the term) going back a few decades. But your point stands.
Unfortunately letters before 1977 weren't available online so I wasn't able to search.
It also helps that I've been to several cities with an actual moat so this word is familiar to me.
I did not mean that it was literally invented a short while ago - a few months ago I had to look up what it means though (not native English).
I wonder how badly this quant affects the output on DeepSeek?
nah. it moat is CUDA and millions of devs using CUDA aka the ecosystem
So far it seems that the best investment is in RAM producers. Unlike compute the ram requirements seem to be stubborn.
With NVDA, you get tools to deploy at scale, maximize utilization, debug errors and perf issues, share HW between workflows, etc. These things are not cheap to develop.
Oh wait, it takes years to do all that and in the meantime you're wasting energy on not staying at the forefront of a hot tech trend.
The higher performing chips, with one less interconnect, is going to give you significantly higher t/s.
Even if you have no interest at all in stock market shorting strategies there is plenty of meaty technical content in here, including some of the clearest summaries I've seen anywhere of the interesting ideas from the DeepSeek v3 and R1 papers.
I remember being surprised at first because I thought it would feel like a wall of text. But it was such a good read and I felt I gained so much.
1: https://youtubetranscriptoptimizer.com/blog/02_what_i_learne...
You may get more milage from excellent writing on a yourname.com. This is a piece that sells you not this product, plus it feels more timeless. In 2050 someone my point to this post. Better if it were on your own name.
I was under the impression that this was not how MoE models work. They are not a collection of independent models, but instead a way of routing to a subset of active parameters at each layer. There is no "expert" that is loaded or unloaded per question. All of the weights are loaded in VRAM, its just a matter of which are actually loaded to the registers for calculation. As far as I could tell from the Deepseek v3/v2 papers, their MoE approach follows this instead of being an explicit collection of experts. If thats the case, theres no VRAM saving to be had using an MOE nor an ability to extract the weights of the expert to run locally (aside from distillation or similar).
If there is someone more versed on the construction of MoE architectures I would love some help understanding what I missed here.
It doesn’t reduce memory usage, as each subsequent token might require different expert buy it reduces per token compute/bandwidth usage. If you place experts in different GPUs, and run batched inference you would see these benefits.
I could be very wrong on how experts work across layers though, I have only done a naive reading on it so far.
I suppose you could look at what part of each layer was most likely to trigger together and segregate those by GPU though
Yes, I think that's what they describe in section 3.4 of the V3 paper. Section 2.1.2 talks about "token-to-expert affinity". I think there's a layer which calculates these affinities (between a token and an expert) and then sends the computation to the GPUs with the right experts.This doesn't sound like it would work if you're running just one chat, as you need all the experts loaded at once if you want to avoid spending lots of time loading and unloading models. But at scale with batches of requests it should work. There's some discussion of this in 2.1.2 but it's beyond my current ability to comprehend!
Let's say you have an inference batch of 128 chats, at layer `i` you take the hidden states, compute their routing, scatter them along with the KV for those layers among GPUs (each one handling different experts), the attention and FF happens on these GPUs (as model params are there) and they get gathered again.
You might be able to avoid the gather by performing the routing on each of the GPUs, but I'm generally guessing here.
If you look deeper, many of these are only applicable to training (we already do FP8 for inference, MTP is to improve training convergence, and DualPipe is to overlapping communication / compute mostly for training purpose too). The efficiency improvement on inference IMHO is overblown.
we already do FP8 for inference
Yes but, for a given size of model, Deepseek claims that a model trained with FP8 will work better than a model quantized to FP8. If that's true then, for a given quality, a native FP8 model will be smaller, and have cheaper inference. If you place experts in different GPUs
Right, this is described in the Deepseek V3 paper (section 3.4 on pages 18-20).I think the title does the article injustice, or maybe it’s too long for people to read to appreciate it (eg the deepseek stuff can be an article within itself).
Whatever the ones with longer attention span will benefit from this read.
Thanks for summarising this up!
Keep it up!
DELETED: If you don't use MTP for drafting, and use MTP to skip generations, sure. But you also need to evaluate your use case to make sure you don't get penalized for doing that. Their evaluation in the paper don't use MTP for generation.
EDIT: Actually, you cannot use MTP other than drafting because you need to fill in these KV caches. So, during generation, you cannot save your compute with MTP (you save memory bandwidth, but this is more complicated for MoE model due to more activated experts).
> Besides things like the rise of humanoid robots, which I suspect is going to take most people by surprise when they are rapidly able to perform a huge number of tasks that currently require an unskilled (or even skilled) human worker (e.g., doing laundry ...
I've always said that the real test for humanoid AI is folding laundry, because it's an incredibly difficult problem. And I'm not talking about giving a machine clothing piece-by-piece flattened so it just has to fold, I'm talking about saying to a robot "There's a dryer full of clothes. Go fold it into separate piles (e.g. underwear, tops, bottoms) and don't mix the husband's clothes with the wife's". That is, something most humans in the developed world have to do a couple times a week.
I've been following some of the big advances in humanoid robot AI, but the above task still seems miles away given current tech. So is the author's quote just more unsubstantiated hype that I'm constantly bombarded with in the AI space, or have there been advancements recently in robot AI that I'm unaware of?
I'm just very wary of looking at that video and saying "Look! It's 90% of the way there! And think how fast AI advances!", because that critical last 10% can often be harder than the first 90% and then some.
There's a huge gulf between what is shown in that video and what is needed to replace a human doing that task.
- An "income effect". You use the thing more because it's cheaper - new usecases come up
- A "substitution effect." You use other things more because of the savings.
I got into this on labor economics here [1] - you have counterintuitive examples with ATMs actually increasing the number of bank branches for several decades.
[1]: https://singlelunch.com/2019/10/21/the-economic-effects-of-a...
DeepSeek is bullish for the semiconductor industry as a whole. Whether it is for Nvidia remains to be seen. Intel was in Nvidia position in 2007 and they didn't want to trade margins for volumes in the phone market. And there they are today.
Theoretically they could be on top if the paradigm changes to big volume slower and lower margin one. But there may be another winner.
Do AMD chips offer more value than Nvidia chips?
- Wait till NVDA rebounds in price.
- Create an OpenAI "competitor" that is powered by Llama or a similar open weights model.
- Obscure the fact that the company runs on this open tech and make it seem like you've developed your own models, but don't outright lie.
- Release an app and whitepaper (whitepaper looks and sounds technical, but is incredibly light on details, you only need to fool some new-grad stock analysts).
- Pay some shady click farms to get your app to the top of Apples charts (you only need it to be there for like 24 hours tops).
- Collect profits from your NVDA short positions.I don’t think this is what happened with DeepSeek. It seems that they’ve genuinely optimized their model for efficiency and used GPUs properly (tiled FP8 trick and FP8 training). And came out on top.
The impact on the NVIDIA stock is ridiculous. DeepSeek took the advantage of flexible GPU architecture (unlike inflexible hardware acceleration).
With Deepseek this is now the 'Pornhub of AI' moment. Adapt or die.
They understood the Dmca brilliantly so they did bulk cheap content purchases and hid behind the Dmca for all non licensed content which was "uploaded by users". They did bulk purchases of cheap content from some studios but that was just a fraction
Of course their risk of going advertise revenue only was high and in the beginning mostly only cam providers would advertise
Our problem was that we had contracts and close relationships with all the big studios so going the Dmca route would have severed these ties for an unknown risk. In hindsight not creating a company which did abuse the Dmca was the right decision. I am very loyal and it would have felt like cheating
Now it's a different story after the credit card shake down when they had to remove millions of videos and be able to provide 2257 documentation for each video
What actually happened was a better algorithm was created and people are betting against the main game in town for running said algorithm.
If someone came up with a CPU-superior AI that'd be worrying for NVidia.
I appreciate China has censorship, but the US is going that way too (recent “issues” for search terms). Might be different scales now, but I think it’ll happen. I don’t care as much if a Chinese company wins the LLM space than I did last year.
Will NVIDIA be in trouble because of DSR1 ? Interpreting Jevon’s effect, if LLMs are “steam engines” and DSR1 brings 90% efficiency improvement for the same performance, more of it will be deployed. This is not considering the increase due to <think> tokens.
More NVIDIA GPUs will be sold to support growing use cases of more efficient LLMs.
I think two threats are the biggest:
First Apple. TSMC’s largest customer. They are already making their own GPUs for their data centers. If they were to sell these to others they would be a major competitor.
You would have the same GPU stack on your on phone, laptop, pc, and data center. Already big developer mind share. Also useful in a world where LLMs run (in part) on the end user’s local machine (like Apple Intelligence).
Second is China - Huawei, Deepseek etc.
Yes - there will be no GPUs from Huawei in the US in this decade. And the Chinese won’t win in a big massive battle. Rather it is going to be death by a thousand cuts.
Just as what happened with the Huawei Mate 60. It is only sold in China but today Apple is loosing business big time in China.
In the same manner OpenAi and Microsoft will have their business hurt by Deepseek even if Deepseek was completely banned in the west.
Likely we will see news on Chinese AI accelerators this year and I wouldn’t be surprised if we soon saw Chinese hyperscalars offering cheaper GPU cloud compute than the west due to a combination of cheaper energy, labor cost, and sheer scale.
Lastly AMD is no threat to NVIDIA as they are far behind and follow the same path with little way of differentiating themselves.
* People have been training models at <fp32 precision for many years, I did this in 2021 and it was already easy in all the major libraries.
* GPU FLOPs are used for many things besides training the final released model.
* Demand for AI is capacity limited, so it's possible and likely that increasing AI/FLOP would not substantially reduce the price of GPUsPerplexity CEO says he tried to hire an AI researcher from Meta, and was told to ‘come back to me when you have 10,000 H100 GPUs’
See https://www.businessinsider.nl/ceo-says-he-tried-to-hire-an-...
Wondering if we are in a similar position with "trust me bro AGI will be achieved with 10x more GPUs".
> DeepSeek is a tiny Chinese company that reportedly has under 200 employees. The story goes that they started out as a quant trading hedge fund similar to TwoSigma or RenTec, but after Xi Jinping cracked down on that space, they used their math and engineering chops to pivot into AI research.
I guess now we have the answer to the question that countless people have already asked: Where could we be if we figured out how to get most math and physics PhDs to work on things other than picking up pennies in front of steamrollers (a.k.a. HFT) again?
There was a crack down on algorithmic trading, but it didn't had much impact and IMO someone higher up definitely does not want to kill these trading firms.
> I would guess most of the profiteering is from consumers not knowing the last transaction prices?
No, not at all. And I wouldn't even necessarily call it profiteering. Ironically, as a retail investor you even benefit from hedge funds and HFTs being a counterpart to your trades: You get on average better (and worst case as good) execution from PFOF.
Institutional investors (which include pension funds, insurances etc.) are a different story.
Another way of saying this: It's a well-known fact that complicated puzzles with a potentially huge reward attached to them attract the brightest people, so I'm arguing that we should be very conscious of the types of puzzles we implicitly come up with, and consider this an externality to be accounted for.
HFT is, to a large extent, a product of policy, in particular Reg NMS, based on the idea that we need to have many competing exchanges to make our markets more efficient. This has worked well in breaking down some inefficiencies, but has created a whole set of new ones, which are the basis of HFT being possible in the first place.
There are various ideas on whether different ways of investing might be more efficient, but these largely focus on benefits to investors (i.e. less money being "drained away" by HFT). What I'm arguing is that the "draining" might not even be the biggest problem, but rather that the people doing it could instead contribute to equally exciting, non-zero sum games instead.
We definitely want to keep around the the part of HFT that contributes to more efficient resource allocation (an inherently hard problem), but wouldn't it be great if we could avoid the part that only works around the kinks of a particular market structure emergent from a particular piece of regulation?
From my personal experience (undergrad physics, worked as engineer, came to CS & ML because I liked the math), there's a lot of pushback.
- I've been told that the math doesn't matter/you don't need math.
- I've heard very prominent researchers say "fuck theorists"
- I've seen papers routinely rejected for improving training techniques with reviewers say "just tune a large model"
- I see papers that show improvements when conditioning comparisons on compute restraints because "not enough datasets" or "but does it scale" (these questions can always be asked but require exponentially more work)
- I've been told I'm gatekeeping for saying "you don't need math to make good models, but you need it to know why your models are wrong" (yes, this is a reference)
- when pointing out math or statistical errors I'm told it doesn't matter
- and much more.
I've heard this from my advisor, dissertation committee, bosses[1], peers, and others (of course, HN). If my experience is short of being rare, I think it explains the grumpy group[2]. But I'm also not too surprised with how common it is in CS for people to claim that everything is easy or that leet code is proof of competence (as opposed to evidence).I think unfortunately the problem is a bit bigger, but it isn't unsolvable. Really, it is "easily" solvable since it just requires us to make different decisions. Meaning _each and every one of us_ has a direct impact on making this change. Maybe I'm grumpy because I want to see this better world. Maybe I'm grumpy because I know it is possible. Maybe I'm grumpy because it is my job to see problems and try to fix them lol
[0] https://bsky.app/starter-pack/roydanroy.bsky.social/3lba5lii... (not perfect, but there's a high correlation and I don't think that's a coincidence)
[1] Even after _demonstrating_ how my points directly improve the product, more than doubling performance on _customer_ data.
[2] not to mention the way experiments are done, since it is stressed in physicists that empirics is not enough. https://www.youtube.com/watch?v=hV41QEKiMlM
Arguably, the emergence of quant hedge funds and private AI research companies is at least as much a symptom of the dysfunctions of academia (and society's compensation of academics on dimensions monetary and beyond) as it is of the ability of Wall Street and Silicon Valley to treat former scientists better than that.
> Is this in academia?
Yes and no. Industry AI research is currently tightly coupled with academic research. Most of the big papers you see are either directly from the big labs or in partnership. Not even labs like Stanford have sufficient compute to train GPT from scratch (maybe enough for DeepSeek). Here's Fei-Fei Li discussing the issue. Stanford has something like 300 GPUs[1]? And those have to be split across labs.The thing is that there's always a pipeline. Academic does most of the low level research, say TRL[2] 1-4, partnerships happen between 4-6, and industry takes over the rest. (with some wiggleroom on these numbers). Much of ML academic research right now is tuning large models, made by big labs. This isn't low TRL. Additionally, a lot of research is rejected for not out-performing technologies that are already at TRL 5-7. See Mamba for a recent example. You could also point to KANs, which are probably around TRL 3.
> Arguably, the emergence of quant hedge funds and private AI research companies is at least as much a symptom of the dysfunctions of academia
Which is where I, again, both agree and disagree. It is not _just_ a symptom of the dysfunction of academia, but _also_ industry. The reason I pointed out the grumpy researchers is because a lot of these people have been discussing techniques that DeepSeek used, long before they were used. DeepSeek looks like what happens when you set these people free. Which is my argument, that we should do that. Scale Maximalists (also alled "Bitter Lesson Maximalists", but I dislike the term) have been dominating ML research, and DeepSeek shows that scale isn't enough. So will hopefully give the mathy people more weight. But then again, is not the common way monopolies fall is because they become too arrogant and incestuous?So mostly, I agree, I'm just pointing out that there is a bit more subtly and I think we need to recognize that to make progress. There are a lot of physicists and mathy people who like ML and have been doing research in the area but are often pushed out because of the thinking I listed. Though part of the success of the quant industry is recognizing that the strong math and modeling skills of physicists generalize pretty well and you go after people who recognize that an equation that describes a spring isn't only useful for springs, but is useful for anything that oscillates. That understanding of math at that level is very powerful and boy are there a lot of people that want the opportunity to demonstrate this in ML, they just never get similar GPU access.
[0] https://www.ft.com/content/d5f91c27-3be8-454a-bea5-bb8ff2a85...
[1] https://archive.is/20241125132313/https://www.thewrap.com/un...
[2] https://en.wikipedia.org/wiki/Technology_readiness_level
- Competition gets crucial features into cheaper hardware
- Work-arounds for most IP are discovered
- Knowledge finds a way out of the castle
This leads to a "Cambrian explosion" of new devices and software that usually gives rise to some game-changing new ways to use the new technology. I'm not sure where we all thought this somehow wouldn't apply to AI. We've seen the pattern with almost every new technology you can think of. It's just how it works. Only the time it takes for patents to expire changes this... so long as everyone respects the patent.
Juicy. Anyone have a link or context to this? I'd not heard of this reception to NOVA and related.
I really don't understand the rationale of "We can now train GPT 4o for 10% the price, so that will bring demand for GPUs down.". If I can train GPT 4o for 10% the price, and I have a budget of 1B USD, that means I'm now going to use the same budget and train my model for 10x as long (or 10x bigger).
At the same time, a lot of small players that couldn't properly train a model before, because the starting point was simply out of their reach, will now be able to purchase equipment that's capable of something of note, and they will buy even more GPUs.
P.S. Yes, I know that the original quote "I think there is a world market for maybe five computers", was taken out of context.
P.S.S. In this rationale, I'm also operating under the assumption that Deepseek numbers are real. Which, given the track record of Chinese companies, is probably not true.
Competition lowers the value of monopolies.
This is already happening today. Most of the new LLM features announced this year are primarily on-device, using the Neural Engine, and the rest is in Private Cloud Compute, which is also using Apple-trained models, on Apple hardware.
The only features using OpenAI for inference are the ones that announce the content came from ChatGPT.
John Gruber says neither Apple nor OpenAI are paying for that deal: https://daringfireball.net/linked/2024/06/13/gurman-openai-a...
NVDA was overpriced a lot already even without r1, the market is full of air GPUs hiding in the capex of tech giants like MSFT.
If orders are canceled or delivery fails for any reason, NVDA’s EPS would be pulled back to its fundamentally justified level
or if all those air GPUs are produced and delivered in recent years, and the demand keeps rising? well, that will be a crazy world then
it's a finance game, not related with the real world
https://proceedings.neurips.cc/paper/2020/file/747e32ab0fea7...
What's totally unclear is what data they used for this reinforcement learning step. How many math problems of the right difficulty with well-defined labeled answers are available on the internet? (I see about 1,000 historical AIME questions, maybe another factor of 10 from other similar contests). Similarly, they mention LeetCode - it looks like there are around 3000 LeetCode questions online. Curious what others think - maybe the reinforcement learning step requires far less data than I would guess?
Is this right? I thought CoT was a prompting method and are we calling the reasoning models as CoT models?
Their software platforms and CUDA are a very strong moat against everyone else. I don't see any beating them on that front right now.
The problem is that I'm afraid that all that money sloshing inside the company is rotting the culture and that will compromise future development.
- Grifters are filling out positions in many orgs only trying to milk it as much as possible.
- Old employees become complacent with their nice RSU packages Rest & Vest.
NVIDIA used to be extremely nimble and was way fighting way above it's weight class. Prior to Mellanox acquisition only around 10k employees and after another 10k more.If there's a real threat to their position at the top of the AI offerings will they be able to roll up the sleeves and get back to work or will the organizations be unable to move ahead.
Long term I think it's inevitable that China will take over the technology leadership. They have the population and they have the education programs and the skill to do this. At the same time in the old western democracies things are becoming stagnant and I even dare to say that the younger generations are declining. In my native country the educational system has collapsed, over 20% kids that finish elementary school cannot read or write. They can mouth-breath and scroll TikTok though but just barely since their attention span is about the same as gold fish.
The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory.
Fortune 100 companies will still want the biggest toolshed to invent the next paradigm or to be the first to get to AGI.
I think longer-term we'll eat up any slack in efficiency by throwing more inference demands at it -- but the shift is tectonic. It's a cultural thing. People got acclimated to shlepping around morbidly obese node packages and stringing together enormous python libraries - meanwhile the deepseek guys out here carving bits and bytes into bare metal. Back to FP!
Is it even legal to give different prices to different customers?
Exciting times to be living in .
Has a wide-scale model analysis been performed inspecting the parameters and their weights for all popular open / available models yet? The impact and effects of disclosed inbound data and tuning parameters on individual vector tokens will prove highly informative and clarifying.
Such analysis will undoubtedly help semi-literate AI folks level up and bridge any gaps.
also thank you for the incredibly informative article.
They are more like the thing which enabled computers to work with and digest text instead of just code. The fact that they can parrot pretty interesting relationships from the texts they've consumed kind of proofs that they are capable of statistically "understanding" what we're trying to talk with them about, so it's a pretty good interface.
But going back to the really valuable content of the books they've been trained on, they just don't understand it. There's other AI which needs to get created which can really learn the concepts taught in those books instead of just the words and the value of the proximities between them.
To learn that other missing part will require hardware just as uniquely powerful and flexible as what Nvidia has to offer. Those companies now optimizing for inference and LLM training will be good at it and have their market share, but they need to ensure that their entire stack is as capable of Nvidia's stack, if they also want to be part of future developments. I don't know if Tenstorrent or Groq are capable of doing this, but I doubt it.
The one thing from the article that sticks out to me is that the author/people are assuming that deepseek needing 1/45th the amount of hardware means that the other 44/45ths large tech companies have invested were wasteful.
Does software not scale to meet hardware? I don't see this as 44/45ths wasted hardware, but as a free increase in the amount of hardware people have. Software needing less hardware means you can run even _more_ software without spending more money, not that you need less hardware, right? (for the top-end, non-embedded use cases).
---
As an aside, the state of the "AI" industry really freaks me out sometimes. Ignoring any sort of short or long term effects on society, jobs, people, etc, just the sheer amount of money and time invested into this one thing is, insane?
Tons of custom processing chips, interconnects, compilers, algorithms, _press releases!_, etc all for one specific field. It's like someone taking the last decade of advances in computers, software, etc, and shoving it in the space of a year. For comparison, Rust 1.0 is 10 years old - I vividly remember the release. And even then it took years to propagate out as a "thing" that people were interested in and invested significant time into. Meanwhile deepseek releases a new model (complete with a customer-facing product name and chat interface, instead of something boring and technical), and in 5 days it's being replicated (to at least some degree) and copied by competitors. Google, Apple, Microsoft, etc are all making custom chips and investing insane amounts of money into different compilers, programming languages, hardware, and research.
It's just, kind of disquieting? Like everyone involved in AI lives in another world operating at breakneck speed, with billions of dollars involved, and the rest of us are just watching from the sidelines. Most of it (LLMs specifically) is no longer exciting to me. It's like, what's the point of spending time on a non-AI related project? We can spend some time writing a nice API and working on a cool feature or making a UI prettier and that's great, and maybe with a good amount of contributors and solid, sustained effort, we can make a cool project that's useful and people enjoy, and earns money to support people if it's commercial. But then for AI, github repos with shiny well-written readmes pop up overnight, tons of text is being written, thought, effort, and billions of dollars get burned or speculated on in an instant on new things, as soon as the next marketing release is posted.
How can the next advancement in graphics, databases, cryptography, etc compete with the sheer amount of societal attention AI receives?
Where does that leave writing software for the rest of us?
I don't think it's necessarily a coincidence that DeepSeek dropped within a short time frame of the announcement of the AI investment initiative by the Trump administration.
The idea is to get the money from investors who want to earn a return. Lower capex is attractive to investors, and DS drops capex dramatically. It makes Chinese AI talent look like the smart, safe bet. Nothing like DS could happen in China unless the powers-that-be knew about it and got some level of control. I'm also willing to bet that this isn't the best they've got.
They're saying "we can deliver the same capabilities for far less, and we're not going to threaten you with a tariff for not complying".
The premise is simple: Business is warfare. Anything you can do to damage or slow down the market leader gives you more time to get caught up. FUD is a powerful force.
My bias comes from having been the subject of such attacks in my prior tech startup. Our technology was destroying the offerings of the market leading multi-billion-dollar global company that pretty much owned the sector. The natural processes of such a beast caused them not to be able to design their way out of a paper bag. We clearly had an advantage. The problem was that we did not have the deep pockets necessary to flood the market with it and take them out.
What did they do?
The started a FUD campaign.
They went to every single large customer and our resellers (this was a hardware/software product) a month or two before the two main industry tradeshows, and lied to them. They promised that they would show market-leading technology "in just a couple of months" and would add comments like "you might want to put your orders on hold until you see this". We had multi-million dollar orders held for months in anticipation of these product unveilings.
And, sure enough, they would announce the new products with a great marketing push at the next tradeshow. All demos were engineered and manipulated to deceive, all of them. Yet, the incredible power of throwing millions of dollars at this effort delivered what they needed, FUD.
The problem with new products is that it takes months for them to be properly validated. So, if the company that had frozen a $5MM order for our products decides to verify the claims of our competitor, it typically took around four months. In four months, they would discover that the new shiny object was shit and less stellar than what they were told. I other words, we won. Right?
No!
The mega-corp would then reassure them that they iterated vast improvements into the design and those would be presented --I kid you not-- at the next tradeshow. Spending millions of dollars they, at this point, denied us of millions of dollars of revenue for approximately one year. FUD, again.
The next tradeshow came and went and the same cycle repeats...it would take months for customers to realize the emperor had no clothes. It was brutal to be on the receiving end of this without the financial horsepower to be able to break through the FUD. It was a marketing arms race and we were unprepared to win it. In this context, the idea that a better mouse trap always wins is just laughable.
This did not end well. They were not going to survive another FUD cycle. Reality eventually comes into play. Except that, in this case, 2008 happened. The economic implosion caught us in serious financial peril due to the damage done by the FUD campaign. Ultimately, it was not survivable and I had to shut down the company.
It took this mega-corp another five years to finally deliver a product that approximated what we had and another five years after that to match and exceed it. I don't even want to imagine how many hundreds of millions they spent on this.
So, long way of saying: China wants to win. No company in China is independent from government forces. This is, without a doubt, a war for supremacy in the AI world. It is my opinion that, while the technology, as described, seems to make sense, it is highly likely that this is yet another form of a FUD campaign to gain time. If they can deny Nvidia (and others) the orders needed to maintain the current pace, they gain time to execute on a strategy that could give them the advantage.
Time will tell.
Anyway, I hope people here find it interesting to read, and I welcome any debate or discussion about my arguments.
Where do you expect NVDA's forward and current eps to land? What revenue drop off are you expecting in late 2025/2026. Part of my bull case for NVDA, continuing, is it's very reasonable multiple on insane revenue. An leveling off can be expected, but I still feel bullish on it hitting $200+ (5 Trillion market cap? on ~195B revenue for Fiscal year 2026 (calendar 2025) at 33 EPS) based on this years revenue according to their guidance and the guidance of the hyperscalers spending. Finding a sell point is a whole different matter to being actively short. I can see the case to take some profits, hard for me to go short, especially in an inflationary environment (tariffs, electric energy, bullying for lower US interest rates).
The scale of production of Grace Hopper and Blackwell amaze me, 800k units of Blackwell coming out this quarter, is there even production room for AMD to get their chips made? (Looking at the new chip factories in Arizona)
R1 might be nice for reducing llm inferencing costs, unsure about the local llama one's accuracy (couldnt get it to correctly spit out the NFL teams and their associated conferences, kept mixing NFL with Euro Football) but I still want to train YOLO vision models on faster chips like A100's vs T4 (4-5x multiples in speed for me).
Lastly, if the Robot/Autonomous vehicle ML wave hits within the next year, (First drones and cars -> factories -> humanoids) I think this compute demand can sustain NVDA compute demand.
The real mystery is how we power all this within 2 years...
* This is not financial advice and some of my numbers might be a little off, still refining my model and verifying sources and numbers
Perhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources.
I wrote in another thread why DeepSeek should increase demand for chips, not lower.1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Paradox.
2. Even if DeepSeek is 45x more efficient (it is not), models will just become 45x+ bigger. It won’t stay small.
3. To build a moat, OpenAI and American AI companies need to up their datacenter spending even more.
4. DeepSeek's breakthrough is in distilling models. You still need a ton of compute to train the foundational model to distill.
5. DeepSeek's conclusion in their paper says more compute is needed for next break through.
6. DeepSeek's model is trained on GPT4o/Sonnet outputs. Again, this reaffirms the fact that in order to take the next step, you need to continue to train better models. Better models will generate better data for next-gen models.
I think DeepSeek hurts OpenAI/Anthropic/Google/Microsoft. I think DeepSeek helps TSMC/Nvidia.
Combined with the emergence of more efficient inference architectures through chain-of-thought models, the aggregate demand for compute could be significantly lower than current projections assume.
This is misguided. Let's think logically about this.More thinking = smarter models
Faster hardware = more thinking
More/newer Nvidia GPUs, better TSMC nodes = faster hardware
Therefore, you can conclude that Nvidia and TSMC demand should go up because of CoT models. In 2025, CoT models are clearly bottlenecked by not having enough compute.
The economics here are compelling: when DeepSeek can match GPT-4 level performance while charging 95% less for API calls, it suggests either NVIDIA's customers are burning cash unnecessarily or margins must come down dramatically.
Or that in order to build a moat, OpenAI/Anthropic/Google and other laps need to double down on even more compute.Fwiw many of the improvements in Deepseek were already in other 'can run on your personal computer' AI's such as Meta's Llama. Deepseek is actually very similar to Llama in efficiency. People were already running that on home computers with M3's.
A couple of examples; Meta's multi-token prediction was specifically implemented as a huge efficiency improvement that was taken up by Deepseek. REcurrent ADaption (READ) was another big win by Meta that Deepseek utilized. Multi-head Latent Attention is another technique, not pioneered by Meta but used by both Deepseek and Llama.
Anyway Deepseek isn't some independent revolution out of nowhere. It's actually very very similar to the existing state of the art and just bundles a whole lot of efficiency gains in one model. There's no secret sauce here. It's much better than what openAI has but that's because openAI seem to have forgotten 'The Bitter Lesson'. They have been going at things in an extremely brute force way.
Anyway why do i point out that Deepseek is very similar to something like Llama? Because Meta's spending 100's of billions on chips to run it. It's pretty damn efficient, especially compared to openAI but they are still spending billions on datacenter build-outs.
Isn't the point of 'The Bitter Lesson' precisely that in the end, brute force wins, and hand-crafted optimizations like the ones you mention llama and deepseek use are bound to lose in the end?
Any customisations that aren't related to the above are destined to be overtaken by someone that can improve the scaling of compute. OpenAI do not seem to be doing as much to improve the scaling of the compute in software terms (they are doing a lot in hardware terms admitedly). They have models at the top of the charts for various benchmarks right now but it feels like a temporary win from chasing those benchmarks outside of the focus of scaling compute.
Keep in mind that the goal everyone is driving towards is AGI, not simply an incremental improvement over the latest model from Open AI.
When was the last time the US got their lunch ate in technology?
Sputnik might be a bit hyperbolic but after using the model all day and as someone who had been thinking of a pro subscription, it is hard to grasp the ramifications.
There is just no good reference point that I can think of.
Then Ethereum turned off PoW mining, so they looked into other things to do with their GPUs, and started DeepSeek.
Intelligence by definition is not compression, but ability to think and act according to new data, based on experience.
Trully AGI models will work on the this principle, not on best compression of as much data as possible.
We need a new approach.
The tldr; if given inputs and a system that can accurately predict the next sequence you can either compress that data using that prediction (arithmetic coding) or you can take actions based on that prediction to achieve an end goal mapping predictions of new inputs to possible outcomes and then taking the path to a goal (AGI). They boil down to one and the same. So it's weird to have someone state they are not the same when it's widely accepted they absolutely are.