Analysts estimate Nvidia owns 98% of the data center GPU market
extremetech.com
extremetech.com
Sure, we all know that the AMD software needs work, but I think the long bet here is on Lisa. Reminds me of the Mac vs. Windows days. If NVIDIA owns this much of the market, then I'd want to work to decentralize my AI business off a single point of failure.
So, I've been working to launch a business to buy up as many of these chips as I can and make them available as bare metal rentals, even with IPMI access. This is something that hasn't been done before. Usually these sorts of AMD cards end up in super computers (ie: Frontier) and/or only accessible to a few people. Azure is catching on with their recent product, but that is a big cloud... I believe people will still want private clouds.
To start with, we're loading up the MI300x chassis with tons of RAM/nvme, top end AMD CPUs and dual 400G networking. We've also got top end management servers for Ray/k8s/slurm. We are interested in feedback on what people want and can be agile enough to customize our purchases for your needs.
My background is that I built a cluster of 150k AMD GPUs across 7 data centers. Deploying and running a lot of compute is something I've gotten very good at. Feel free to reach out.
I do fully recognize that in order to build the developer flywheel, we need to solve both problems in the long run. For now, there are other great companies that are focused on solving the software side, such as NeuralFabric, EmbeddedLLM, Lamini and MK1. Our intention is to partner and be friends with all of them. There is a lot of room in the picks/shovels ecosystem for all of us.
Dual 9754's are indeed quite nice. ;-)
The GPUs are to train a model. Separate use-case, but we use same hardware for both use-cases so that's why the big cache thing.
I would like you to not by pallets of these chips so that I and others may be able to get our hands on a couple for our own personal projects.
Case in point... try searching ebay for the previous generation... MI250.
If you want access to these chips, you're going to have to rent them. Given that nobody is really renting them currently, at least I'm working to get you access to them at all.
The chip industry runs on relationships.
I did personally sign some documentation (EAR) saying that I wouldn't export the equipment to a whole list of countries and I definitely won't do that. I also won't rent to anyone in those countries either, just to cover my bases. I have high ethics and have no desire to get in any sort of trouble.
That "no consumer hardware support" sounds stupid, but people getting into ML/AI, grad students, etc want to be able to mess around and develop/prototype on local hardware they already own.
You cannot do that with ROCM, because they only support a very, very small number of consumer-level cards - and very oddly, it's a mix of the very high end and low end, nothing in the middle.
AMD is also massively behind hardware-wise with only the current 7xxx series cards that just came out having AI-specific hardware, whereas NVIDIA has had tensor cores for three generations of their cards.
But of course AMD can't do anything right, so they nerfed the GPU's graphics processing cores when they added the AI stuff, making the cards a worse deal from a pure gaming standpoint. The 7000 series cars are basically a tiny bump at best, and in many games worse, than the 6000 series equivalents. Their only advantage is better hardware video encoding support and somewhat better power consumption.
The only difference is that previously you couldn't even rent time on these super high end AIA's. They were reserved for supercomputers only. Even the MI250x sku is government/research only and even I am unable to buy that over the MI250.
So, we need to work on building that developer flywheel. My view is that if you can rent some time on one of these systems, that is a good first step. Nowhere else can you load a model into 192gb of RAM on a top end system.
You might be able to argue that MI300 is competitive, but it is certainly hard to argue it's superior, so we're not yet at the GPU equivalent of the 2017 release of EPYC.
Meanwhile those people then know each other, so as soon as someone credible claims it works, several other credible believe them, try it themselves, and say the same thing. It doesn't really matter if you think it's 10 or 100 or 1.2, it's exponential spread at the speed of the internet.
What matters is not just "the software is working", but also "the software will keep working" (for both new releases and new hardware). Reputation is meaningful for that.
This is why anyone making long-term investments in hardware-related software is wise to write it in a vendor-agnostic way -- even if the software continues to be good you may have other reasons to want to switch hardware in the next generation.
1. Code must be hand optimized, refactored for each architecture while the API stays the same.
2. New APIs, algorithms must be backported and optimized for older hardware.
That's insane amount of testing and performance tuning.
One is, can I just buy this thing and run the existing somebody else's code on it and it will work and be faster than the previous generation? If it isn't, you just don't buy the new hardware until that changes.
The other is, you're going to buy a thousand of them or you're making software for the general public to use and you want to optimize it for each specific generation. But in that case you're redoing the work for each generation regardless and you don't have to be concerned about your current efforts carrying forward because you already know that they won't.
I'd be skeptical that NVidia can maintain that margin, or price point. Had they owned ARM it may have been possible, as new entrants to the chip market could have been locked out.
You are not far off.
Nvidia spends $3,320 to manufacture H100 unit, says Raymond James Financial, Inc. analysis.
If the AI hype is real, then Nvidia revenue is $300 billion in 5 years. Assuming P/E 25 would mean $7.5 trillion valuation. That sounds insane.
If hype is 1/3 true, Nvidia grows 17% per year and is valued $2.5 trillion in five years.
ps. Intel vs. AMD is not completely symmetric competition because AMD is fabless and Intel is not. While AMD gets better margins, Intel can produce volumes to match demand. AMD competes directly with Apple, Nvidia, ARM,.. for TSMC fab capacity (I know Intel also buys manufacturing from TSMC).
25 * $300B (revenue) = $7.5T, yes.
But that's using a P/S multiple, not P/E. Your math makes sense if you assume Nvidia will have 25x P/S... which historically is very high, and unlikely to be the "correct" multiple after all of that growth has already occurred. Unless you expect them to continue growing at same CAGR in perpetuity.
Also, very unlikely they will maintain that kind of market share if they reach those valuations. A relatively simple software solution to improve integration with AMD/Intel/TPU offerings that is worth trillions is a bit of a no brainer. It wasn't a financial prerogative until recently, so yes, all the competitor software here sucks, today.
Compute is fungible, despite what many say here. The cost differential between implementing good software support for AMD and paying Nvidia ~98% gross margin makes it kind of obvious.
A bull case could still see NVDA at a few trillion 5-10 years out... very unlikely it would go much beyond that. I fail to see any remotely rationally made, objective, case for it. At least not one that warrants investing in Nvidia today, at the current valuation, versus other opportunities
It can still be a multi-bagger in the short-term due to what looks like dotcom 2.0 enthusiasm.
Nvidia has committed to CUDA and created an ecosystem. This has been a long term commitment. Meanwhile, Intel (new libaries/paradigms all the time) and AMD (software/drivers/libraries are an afterthought) have struggled in this arena.
Maybe some regulatory body could declare this an obvious 'monopoly' and at least force the 'language spoken' to be free for anyone to implement without patent worries?
Company profits are not the top priority, if it stuffers so that actually important values are preserved, there's no problem with that at all.
Misusing monopoly power is illegal.
Nvidia has monopoly in the market, but unless they use their position illeagally, that's OK. GPU market is competitive and Nvidia has put in the work others like AMD or Intel never bothered to do.
(see: Lina Khan’s management of the FTC, or “New Brandeis movement”)
Lina Khan and others are fighting good fight against misuse of monopoly powers, two sided markets, anticompetitive effects of platform-based business models and so on.
GPU market is not natural monopoly, nor is the market structured as monopoly. Nvidia just has a competitive advantage.
OK? I sincerely disagree.
This lead won't last forever. Everyone in the industry has a vested interest in this collapsing, and there are ongoing assaults from all angles.
George Hotz' new company is working to undo this, and there are lots of others. All potential AMD acquisition targets, too.
There are lots of little OpenCL-like projects, and consensus is starting to form.
Then there's the growing TPU market...
Give it two years.
Is that before or after George Hotz solves homelessness and drug addiction and saves democracy (https://news.ycombinator.com/item?id=39206959)
Absolutely nothing will change about the context in two years, other than Nvidia will be far stronger at that point with $50+ billion per year in operating income (formally joining the hyper profitable US tech giants).
It's going to be extraordinarily expensive to build in the GPU data center space at scale. Nvidia will have the money to do it and most of their smaller competitors will not (that includes AMD, which doesn't generate anywhere near enough cash to outlay billions of dollars in a single high risk capital investment pointed at the GPU data center space; AMD's operating income over the last four quarters is negative, and they have a mere $3b in cash). Nvidia will use their new cash spigot to build infrastructure moats in the space. Large governments, Microsoft, Apple, Meta, Google, TSMC are the only entities capable of keeping up on spending.
It isn't too late and I'm betting it will be closer to 5 than 10.
> It's going to be extraordinarily expensive to build in the GPU data center space at scale.
It already is expensive. But it isn't just NVIDIA doing it. CoreWeave is the largest and certainly backed by NVIDIA, but they are a separate business. AMD doesn't have to do it on their own either.
It's also weird that you're looking at operating income, which already has R&D expenses subtracted out of it. A company spending more of its revenue on R&D will have a lower operating income, but that hardly implies they can't spend on R&D -- because they are.
And customers hate moats. In many cases they get stuck with them, but most often this is individual consumers with no resources to do anything about it. Data center customers are large institutions, often with their own R&D budgets. Three of the companies that each have a larger market cap than Nvidia are Amazon, Google and Microsoft, the three largest cloud providers. Is any of them interested in letting Nvidia have a moat?
C++, Fortran, anything else with a compiler toolchain against PTX, graphical debugging and IDE integration, and a library ecosystem.
Khronos was sceptical that Fortran was even meaningful, and SPIR only came up after NVIDIA was approaching the finish line.
Intel and AMD only have themselves to blame.
https://www.macrotrends.net/stocks/charts/AMD/amd/market-cap
Now they have to take the money and use it to fix their software. It's pretty obvious why they didn't do that in 2008.
AI has been the first technology to really kick them in the pants to get off their asses and really invest in that aspect. There are a ton of job openings around AI software now within AMD. I'm seeing new ones get posted almost daily on LI. It'll take time, but they will fix it.
The world's largest computer, Frontier at oak ridge national laboratory, runs AMD GPUs. AMD is undisputably the #2 in the GPU space.