Nvidia's CUDA Monopoly
matt-rickard.com
matt-rickard.com
Nvidia is pushing upmarket, focusing on data center products it can charge huge amounts for.
As Nvidia pushes upmarket, the traditional computing market forces will come into play. These forces have played out for 50 years.
Users will seek cheap and available GPUs and they’ll find a way to get the job done with them.
At the moment, people say they can’t use retail GPUs because of RAM constraints.
This will change however. Through necessity, software will be developed that gets the job done on consumer grade GPUs. It might be open source. It might be AMD or Intel software.
This has always been the way with computers…. innovation at the low end will foil the plans of IBM to control the entire market. Oops did I say IBM? I meant Nvidia.
Nvidia has the chance to own everything long term, but monopolists can’t help but become greedy and become their own worst enemy. Nvidia is milking its customers for huge, huge profits. The customers will find another way, this Nvidia will have indirectly created its true competition.
If Intel and AMD want to defeat Nvidia they need to not play Nvidias game of going for the high end and turning its nose up at the low end. AMD and Intel need to produce the lowest cost, most powerful GPUs they can. They also need to attack Nvidia where it hurts. Nvidia has artificial constraints on its GPUs. Through drivers it prevents certain uses so that customers are forced to use high end data center GPUs. Where ever Nvidia has artificial constraints, AMD and Intel need to NOT have those constraints.
Where Nvidia is closed source, Intel and AMD need to be open source.
Nvidias dominance won’t end, but it’s monopoly will, and viable competition will form simply because Nvidia is so anti consumer, and in the computing game consumers are very resourceful.
What’s needed is a bunch of smart programmers who know how to squeeze usable AI out of a 24 gigabyte AMD 7900 XTX.
Tonight I'm going to play around with LLMs. I'm probably going to go to bed early because my AMD graphics card running ROCm (on an unsupported system and a probably unsupported graphics card) will eventually cause a kernel panic.
In some sense that is unforgivable. In other senses, they are only 1-2 bugs away from being perfectly good enough and competitive with Nvidia for me. That isn't much of a moat. It usually takes an hour or so for the drivers to collapse and for those hours everything works great.
AMD also said this 24 months ago.
We'll see how smart the market and the monopolist are this time around, yes, in the next 12-24 months... and maybe longer.
I was specifically referring to AMD's investments in a unified computing stack (their equivalent of CUDA) which were supposed to bring them to computational parity in the AI/ML.
https://www.extremetech.com/computing/here-we-go-again-ai-de...
Retail gaming GPUs are available and they’re MUCH cheaper than cloud GPU.
It’s happening now.
AI software needs to make a choice too, does it only run in Nvidia, or does it make itself compatible with the mass market.
For 50 years software developers have chosen to make their software work on cheap mass market devices.
However, some ML models (such as LLMs) demand a lot of vram to train. Simultaneously, a lot of gamers have been unimpressed by the latest generations of gaming GPUs due to the limited vram.
nvidia's GTX-1070 released in 2016 with 8 GB of vram and an MSRP of $379. Then the RTX-2060 Super, with 8 GB of vram and an MSRP of $399. Then the RTX-3060 Ti, with 8 GB of vram and an MSRP of $399. Then the RTX-4060 Ti with - you guessed it - 8 GB of vram and an MSRP of $399.
Some people think nvidia is deliberately being miserly with vram on gaming cards, to try and force ML users onto the $$$$$ data centre GPUs.
I suspect this is what andrewstuart means by "Nvidia is pushing upmarket" - that they're deliberately letting their gaming products languish, in pursuit of the data centre market.
4GB (from personal experience, 2GB from what I’ve heard) is enough to run stable diffusion 1.x/2.x with the major current consumer UIs and the optimizations they make under the hood. Not sure about SDXL. The original first-party inferencing code takes more.
> Some people think nvidia is deliberately being miserly with vram on gaming cards, to try and force ML users onto the $$$$$ data centre GPUs.
And, sure, as long as there is no real competition, that kind fo segmentation makes sense. OTOH, if there is competition and desktop ML demand, it will make progressively less sense. So, the question seems to be, will there be competition that matters?
Who's making that choice? It sounds like a lot of people want to make this stuff run well on iPhone and Android and AMD and Mac, but they just don't have the API control or insider help. This isn't a thing where Open Source developers step up to the task and fix everything with a magic wand. This is a situation where the entire industry converges on a compute standard, or Nvidia continues to dominate. They are betting on Apple, AMD and Intel being stuck in an eternal grudge-match, and their residuals won't stop rolling in.
The real question is much less hopeful; can the industry put aside it's differences to make a net positive experience for consumers?
...probably not. None of the successful companies do, anyways.
I say this as a person who's diving right into the NVidia moat. "Well everyone gets NVIdia GPUs so that's what I should get".
this isn't in the hands of the industry as a whole, though. The question is: Does AMD think they can earn a better profit down the road than if they invest in/focus on something else (e.g. maintaining their lead over intel on x86, continue being dominant in the console gaming GPU sector, etc...).
So what? “The industry as a whole” isn’t an actor that makes decisions.
The first thing that's needed is for AMD to wake up and invest into its tooling stack.
What's AMD tools look like and what would it take to catch up with CUDA
and I guess it does not depend on AMD alone, the cloud providers trying to get alternatives to Nvidia could be investing into this too
Maybe at one point they wanted to beat CUDA, but they are pretty happy with the feature segmentation now.
At least by revenue, consoles are now the leader in gaming anyway, the PC market is shrinking - the old dudes crowd is aging out of hi-perf gaming (due to work and having kids), and the young crowd is either on hassle-free consoles or mobile). And a lot of the "gaming" GPU marketshare of the last few years were coin miners.
If AMD wants to be known for something else than tying their future to two large companies who can always hop back to NVIDIA, they absolutely have to step up their compute game to make sure they at least have the ability to pivot, should either of their deals with Microsoft and Sony fall apart. NVIDIA is the best example, it's a miracle Soldergate didn't sink them entirely a decade ago.
[1] https://www.gamesindustry.biz/report-pc-and-console-global-g...
What would be the motivation of said programmers? Who will pay them?
Whether the Valve situation strikes twice though isn't guaranteed. We already see all the other PC software stores continue to crowd around Windows providing apologies for its actions. If an AI firm can see past the next quarter, it could have huge ramifications.
(cool bug facts: for the Q1 numbers, AMD's gaming division actually had a higher operating margin than NVIDIA as a whole including enterprise sales/etc! 17.9% vs 15.7%)
https://www.macrotrends.net/stocks/charts/NVDA/nvidia/operat...
https://ir.amd.com/news-events/press-releases/detail/1146/am...
The new generations of goods simply are that expensive to produce and support - TSMC N4 is something like 3-4x the cost per wafer of Samsung 8nm, even with the smaller die sizes and smaller memory buses these products are actually some thing on the order of twice as expensive as Ampere to produce. Tapeout/validation costs have continued to soar at an equal rate, and this is simply a matter of physics, not something TSMC controls, therefore not something that can be changed by pushing profits around on a sheet.
https://www.tomshardware.com/news/tsmc-will-charge-20000-per...
https://semiengineering.com/how-much-will-that-chip-cost/
MCM doesn't really affect this either - this is about how many mm2 of silicon you use, not how well it yields. Even if you yield at 100%, 4x250mm2 chiplets is still 1000mm2 of silicon. And the cost of that silicon is increasing 50-75% for every node family you shrink. Memory and PCIe PHYs do not shrink, and cache has only shrunk modestly (~30%) at 5nm and will not shrink at 3nm at all. So at the low end you are getting an additional crunch where the design area tends to be dominated by these large fixed, unshrinkable areas.
This is the fundamental reason that has followed NVIDIA on pricing for years now with products like 5700XT, 6800XT, Vega, Fury X, etc. TSMC N7 was already around twice as expensive per wafer as Samsung 8nm or TSMC 16FF/14FF. NVIDIA was using cheap nodes and AMD was being forced to use expensive leading-edge nodes just to remain competitive (and despite failing to take a commanding lead). They didn't follow because their cost-of-goods was higher - which is also the reason AMD cut memory bus width in RDNA2, along with PCIe width and things like video encoders on certain products. They were on the leading nodes, so they took the hit to area overhead sooner, and started looking for solutions sooner. It's not oligopoly, consumers just haven't adapted to the reality of cost-of-goods going up.
AMD being willing to cut deeply on inventory at the end of the generation is one thing, but you also have to remember that the Gaming Division (RTG+consoles) is barely in the green right now and that's even with AMD signing a big deal for next-gen consoles that is more or less straight profit for them. They're losing money on RDNA2 hand-over-fist right now, this is a clearance sale at the end of the generation, not a sustainable business model.
When they can undercut they have undercut - like 7900XT/XTX. The 4080 is fundamentally unsaleable at its current price, and 4070 Ti isn't exactly great either, and when AMD saw the sales numbers they cut the prices. That's the counterexample to oligoploy - they are willing to do it when they can. When they match pricing, or do a token undercut, it's usually because they can't afford to do drastically less than NVIDIA themselves, because they're affected by the same costs. For example Navi 32 (7700XT/7800XT) is going to be pretty unexciting because MCM imposes a big area/performance overhead and they simply can't go way cheaper than NVIDIA, despite the 4060 Ti also being an awful price.
As such, the premise of your post is fundamentally false. AMD generally costs about the same as NVIDIA, and that's part of the reason they've failed to take marketshare over the last 10 years. It's generally a less attractive overall package with less features, other than sometimes having more VRAM (but not always, Vega had less than 1080 Ti, Fury X had less than 980 Ti, 5700XT had the same as 2070/2070S and less than 1080 Ti, etc). And this is because AMD is fundamentally affected by the same economics in the market as NVIDIA. And Intel will be too, when they stop running loss-leader products to build marketshare.
Consumers don't like it but this is the reality of what these products cost to design and manufacture today. And you don't have to buy it, and if those product segments don't turn a profit they will be discontinued, like the $100 and $150 segments before this (where are the Radeon HD 7750s and 7850s of yesteryear?). That's how capitalism works, the operating margins have already been reduced for both companies, they're not going to operate product segments at a loss just so gamers can have cheap toys. Even if they do it for a while, they're not going to produce followups to those money-losing products.
Nor is NVIDIA leaving the gaming market either. They'll continue to produce for the segments that it makes sense to produce for. The midrange (which is now $600+) and high-end ($1000+) will both continue to see good gains, because they're less affected by the factors killing the low end. And MCM will actually be incredibly great in the high end - imagine two or four AD102-sized chiplets working together. But it's going to cost a ton - probably the high-end will range up to $4k-8k within a decade or two.
The low end will have to live with what it's got. AMD holding the 7600/7500 back on N6 (7nm family) is the wave of the future, and it seems like NVIDIA probably should have done the same thing with the 4060. A 3060 Ti 16GB for $349 or $329 would probably have been a more attractive product than a 4060 8GB at $299 on 4nm. Maybe give it the updated OFA block so it can use the new features, and call it a day.
If $300 is your budget for a GPU, buy a console - the APU design is simply a more efficient way to go. A $300 dGPU is about 90% of the hardware that needs to go into a console, and if you just bolt on some CPU cores and an SSD you're basically there. The manufacturing/testing/shipping costs don't make sense for a low-end product (and $300 is now entry-level) to be modularized like this anymore, integration brings down costs. The ATX design is a wildly inefficient way to build a PC, and consoles eschew it for a reason. A Steam Console could bring costs down a lot, but it still will involve soldered memory and other compromises the PC market doesn't like (but will be necessary for GDDR6/7 signal integrity etc). Sockets suck, they are awful electrically and ruin signal integrity. The future is either console-style GDDR or Apple-style stacked LPDDR.
I wouldn't go as far as saying Nvidia is making less just because their net income is less, couldn't they still be huge but consumed by even greater investment into future?
CPU integrated GPU is what's killing the low end. You don't need any more than that for desktop/light gaming.
If iGPUs can get to punching in the same weight class as the RTX 3060s, that will be a paradigm shift for gamers and most household graphics users.
Of course, if Intel puts more emphasis on ARC iGPUs I wouldn't be entirely surprised to see mobos come out with GDDR VRAM soldered on. Some people will definitely still screech, but they aren't important.
On discrete GPUs memory is placed in a circle around the GPU chip, each memory module connected with individual wires.
This will never happen; a discrete card will always be faster. The interesting question rather is: will or when will iGPUs become fast enough for most purposes that currently still require a discrete GPU?
And a Ferrari will always be faster than a Ford. But, for most people, the Ford is good enough to do what they want. The largest gaming segment (consoles) is already powered by APUs. Discrete GPU sales have been in a steady decline since 2009 [1].
[1] https://siliconangle.com/2022/12/29/sales-computer-graphics-...
In the interest of big corporations is to have profits as low as possible, so that they don't pay excessive amounts of Corporation Tax.
All they wanted is Big equals Evil. And the classic underdog story.
But it is nice to see there is at least one more person on HN understanding hardware.
https://github.com/RadeonOpenCompute/ROCm/graphs/code-freque...
https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...
Then maybe, most researchers will chose anything other than CUDA.
I'm not sure about the state of intel software stack or hardware. Amd hasn't reached hardware parody either with some of the specialized execution units that are available across all most of nvidia's products. Perhaps that's low hanging fruit for AMD?
I'm looking forward to the day that somebody fully upturns this market with products that are available to end the consumers.
The reason they still do so is because the majority of voters aren’t too bothered by it.
voters dont have control of their representatives either way
Maybe your reading your own thoughts into someone else?
We, the consumers, don't get a damn word about what we want. Unless the free market says it's a bad thing, you're stuck enjoying whatever the corps decide the fight-of-the-week is.
Though I don’t believe this line of reasoning applies to CUDA/GPUs (yet).
that was my only point, they put monopoly in quotes to attempt to defend Nvidia from regulatory scrutiny
they havent done the thing that causes regulatory scrutiny while it still being accurate about what they have achieved
at larger levels there are typically business units within a larger organization that might have a monopoly in that sector
if everyone chooses your phone brand out of social ramifications and you dont do anything to accelerate competitor’s exiting the market, is what it is
I was really encouraged to see this the other day: https://news.ycombinator.com/item?id=36968273
Because the AMD GPU FFT library isn't feature complete compared to cuFFT or even FFTW.
Back then, everyone thought Windows NT had “won”, and it was game over for every other operating system.
Same thing here. Nvidia might look strong now but all it takes is an open source project that competes effectively to change everything.
For practical purposes it is. Bur it's not Windows NT that killed it, it's Linux.
The actually reality is POSIX flavoured OSes are most widely used operating systems in the world, for various levels of what means to be POSIX, and to what level they expose it.
"Transcending POSIX: The End of an Era?"
https://www.usenix.org/publications/loginonline/transcending...
Additionly Darwin evolved from a mix of BSD and MACH.
NeXTSTEP drivers were written in Objective-C, and a C++ subset on Darwin.
In any case you won't be running CLI applications on iOS, rather applications that depend on a mix of Objective-C and nowadays also Swift, hardly UNIX.
Using BSD sockets on iOS doesn't support all the iOS networking capabilities, for example.
NeXT approach, later adopted by Apple, was similar to Windows NT POSIX in spirit.
The UNIX compatibility is there to help bring applications into the platform, and take place in US government contracts, not to make it easier to port applications elsewhere.
There are even some recordings from meetings at NeXT, with Steve Jobs discussing how NeXT should position itself on the workstation market.
Nvidia made their fortune riding 2 consecutive hype waves - crypto and LLMs.
That's a pretty reductive way of looking at things. Nvidia got to where they are today because they made the "different better AI model" by spending billions of dollars in AI research when other companies were fighting over scraps. They were the ones investing in a high-level GPGPU programming model when Apple and AMD were busy stabbing each other in the back over OpenCL. Just look at their list of papers and compare it to the contributions of your favorite FAANG member: https://research.nvidia.com/research-area/machine-learning-a...
Nvidia owns this market because Apple, AMD and Intel have a grudge match that is apparently more important than cross-platform harmony.
I attended one Khronos webminar where the panel was surprised that anyone would ask about Fortran support, and showed total lack of knowledge about PGI compilers.
Additionally with very crude tooling.
Everyone is out for blood in this market. Going after Khronos for not addressing Fortran FFI is a cart before the horse argument, from where I'm standing.
I see this:
https://github.com/openai/triton/issues/1073
But it's not clear to me if we will see AMD GPUs as first class citizens for pytorch in the future?
...Yeah.
Nvidia seems well aware that the "There is existing code many users want to run written in CUDA, and you can only run CUDA code on an Nvidia part" situation is their competitive advantage.
ATI/AMD's failure to settle on a stable GPGPU toolchain (CTM/THIN/Brook+, Stream, ROCm with OpenCL, HIP, and perpetually broken CUDA compat...) and OpenCL's ugly boilerplate gave them an opportunity to get that core set of lock-in software, and they're not giving it up without a fight.
Which the GP computing world broke out of with the advent of multi vendor ISA compatibility in commodity CPUs (x86, later others) and open source compilers (GCC, later others). Before that the sw stacks were buggy, vendor locked, expensive, were gatekeeping programming langugaes development, very similarly as now in GPUs.
Not to mention extension spaghetti with proprietary features anyway.
Sometimes it sounds more like gushing than anything else, always the best, consistently sold out, etc etc.
Nobody stops a hardware company to hire software developers. The other way around would be harder.
They tend to lack appreciation/understanding of software in the first place (hardware-first thinking) leading to underinvestments.
It's also hard for them to identify great software people - at all levels, starting from the CEO/board.
[0] https://github.com/RadeonOpenCompute/ROCm/issues/666#issueco...
[1] https://github.com/ggerganov/llama.cpp/discussions/915#discu...
They decided to fight Nvidia on the gaming segment and let them win productivity. It seemed like a good strategy a few years ago when GPU were unobtainium, not so much today.
It is only one of the biggest selling consoles in the history of game consoles, breaking the record of GameBoy and PS 4, not a big deal.
Sony/Microsoft literally lose money on each console sold hoping to recoup in subscription and store costs.
The margins in question here are the margins for Nvidia/AMD, not Sony/Microsoft/Nintendo. It would not surprise me at all if AMD makes better margins on the Playstation than Nvidia does on the Switch.
You're asking why a smaller, less funded team, doesn't throw massive resources into following an encumbent with multi-year head start and questionable ability to extract a big profit... vs. spending resources into entering disjoint markets where they can win (e.g. being the vendor of choice of almost every single console on the market - PS5, Xbox, Deck, etc.).
Rocm seems perpetually imminent…