Nvidia H100 and A100 GPUs – comparing available capacity at GPU cloud providers
llm-utils.org
llm-utils.org
Somewhat related and hopefully helpful: My experience so far using the H100 PCIes to train ViTs:
Pros: A little over 2x performance compared to A100 40GB. 80GB by default. fp8 support, which I haven't played with but supposedly is another 2x performance win for LLMs. For datacenters and local workstations it's twice as power efficient as the A100s. Supposedly they have better multi-GPU and multi-node bandwidth, but my workload doesn't stress that and I can only rent 1xH100s at the moment.
Cons: I had trouble using them with anything but the most recent nVidia docker containers. The somewhat official PyTorch containers didn't work, nor did brewing my own. Luckily the nVidia containers have worked fine so far. I just don't like that they use nightly PyTorch. In addition to that, because of their increased performance, they really start to push the limits on feeding data fast enough to them with existing system configurations. I was CPU-limited on the LambdaLabs 1xH100 machines because of this.
Overall the pricing has worked out equal for my use-case, but I'm sure fp8 would make them more affordable. If fp8 were available out-of-the-box on PyTorch I'd play with it, but it's only available right now from some nVidia codebase specific to LLMs.
Even at equal pricing, having twice the power per GPU and per node is a big win. That increases experiment iteration across the board.
Side note: If I recall correctly, nVidia is heavily differentiating the H100 products by their interface this go around, which I find quite odd. The SXM version of the cards are supposed to be something like twice as beefy as the PCIE version? Not sure why they're doing that; gonna make comparing rentable instances all the more difficult if you overlook that little detail. A100 had a little bit of this, but the difference was never much in practice except between specifically the A100 40 GB PCIe and the 80 GB SXM, which was something like 10% faster.
Anyone else having fun with the new toy our overlords have allowed us to play with?
I can't help but feel it's another cloud/enterprise cash grab. The SXM baseboards are way more expensive than server motherboards and NVIDIA makes them.
SXM baseboards are made by more than Nvidia, Dell, HP, and Supermicro all have their own designs.
(Disclaimer - I work for MS Azure but have no internal knowledge about the costs/designs/capex whatever of these systems)
Some automatically drop those limits at a certain age, others require you talk to them.
We have users that have a workstation with a 3090, but the larger amount of memory is what they are after.
I have a 2x3090 rig as my local machine, which has been useful for early experimentation, but I'm at the stage now where my runs are at 200 million samples which would take ages on that rig. 8xA100 can do it in tens of hours, which allows me to iterate faster.
An 8x 3090 or 4090 machine locally would be great, but is a huge hassle to build. Last I looked into it there really wasn't a lot of knowledge available online on how to even do it. I did find an EPYC server motherboard and such that I could theoretically use, but couldn't find a great source for a >3,200 Watt server power supply. Everything in that domain is geared towards either B2LargeCorp, or B2ServerBuilders2SmallBusinesses.
I could of course buy an 8xA100 rig no problem for some ungodly amount of money, but as noted above that's not appropriate for this project.
In both cases, I now have to figure out what to do with 3,200kW of heat output in my office which I'm trying to avoid turning into sauna. Or co-lo it for more money out of pocket.
So renting off the cloud has worked well enough at this scale.
So how much did it cost?
I'll stick with 2x 3090s for now, but maybe I'll go for an 8x monster in a few years when the hardware shakes out a bit if there is a good reason to do so.
Power draw. The 4090 pulls 450 watts, and is destroying connectors. H100 SXM uses 700 watts.
If the PCI consortium wants NVIDIA to use their standard for top end products, they need to make a better connector.
I was talking about the SXM interface instead of PCIe
Also, that’s good feedback on the CPUs bottlenecking. I’ll let our HPC hardware team know about this.
We are also looking into GPU direct storage to help resolve this:
Each H100 is paired with 32 3.25GHz CPUs
Unless something has changed in the last few days, I was literally just using an H100 on LambdaLabs without having to speak to a sales rep…
If there was real competition -like two or more other suppliers on par in terms of capability - then artificially constraining devices just to plump up prices would not be a thing.
Citation?
To the contrary, millions of consumer-level Nvidia customers have access to datacenter-grade HPC APIs because of their vertical integration. Nvidia's "monopoly hold" on GPGPU compute exists because the other competitors (eg. AMD and Apple) completely abandoned OpenCL. When the time came to build a successor, neither company ante'd up. So now we're here.
CUDA is not a monopoly. If Apple or Microsoft wanted, they could start translating CUDA calls into native instructions for their own hardware. They don't though, because it would be an investment that doesn't make sense for their customers, costs tens of millions of dollars, and wouldn't meaningfully hurt Nvidia unless it was Open Source.
Nvidia is certainly hostile to Open Source and not the kindest hardware vendor to boot, but that alone does not suffice a monopoly.
You're right.
CUDA is what does.
It might be an argument that, to the extent that that is the sole basis for the monopoly, the monopoly is unlikely to be a long-term stable condition, but its not a counterargument to it existing.
No one said it was. Having a monopoly isn’t illegal in the US at all (leveraging it in certain ways is.)
The claim was that (1) NVidia has a monopoly, and (2) the effect of that monopoly has been consumer devices getting worse for this use in specific, well-defined ways. Legality of NVidia’s actions and fairness of how their market position arose are not particularly relevant.
AMD doesn't care and it's difficult to blame that on NVIDIA.
Also, a consumer-grade GPU can used for neural net training at the researcher level but large corporate use requires H100/A100 and that is what's getting traction.
For Apple and AMD, that's not really a problem. Both of them drive considerable (40%+) margins on their products and can afford to drive things closer to the wire.
I also think more competition here would be good (and I do love lower prices) but Nvidia charges more here because they know they can. It's value-based marketing that works, because their software APIs aren't vaporware.
> large corporate use requires H100/A100 and that is what's getting traction.
I guess... you really need a strict definition of "requires" for that to hold true. For every non-"competing with ChatGPT" application, you could probably train and deploy with consumer-grade cards. You're technically right here though, and it invites the conversation around what actually constitutes abusive market positioning. Nvidia's actions here really aren't much different than AMD and Intel separating their datacenter and PC product lines. It's a risky move from a "keeping both users happy" standpoint, but hardly anticompetitive.
They could that - but the reason they command these margins is exactly because they don't do that. I think do something like for fairly some investment but producing products that would compete with Nvidia would require a significant percentage amount of capital for any company - those dealing with tens of billions of dollar chunks expect above commodity revenues.
Then I guess we're back around to square 1 again. Is this fair, or anticompetitive behavior?
But, putting on my economist hat, not all market-structures naturally generate large-scale competition in the fashion of white box PC clones. Some market structures are naturally monopolies (energy), some are naturally oligopolies (automobiles) and some naturally have a dominant player plus marginal players arrayed around them.
There's just no easy solution to this. That said, it's not like we don't GPUs of unprecedented power available at a variety of price levels.
But these specific critisms don't really ring true. Lower VRAM bandwidth lets them use lower binned VRAM and there isn't really a need for more RAM than the 24G in the 4090 in gaming.
The naming screw up they did with the 4080 was dumb and fortunately corrected quickly. But it doesn't seem related to the OPs point.
For example. An '8 core/16GB RAM/3080 Ti 12GB' is 50% more expensive on salad than on Vast.
In fact, I can get more than double the resources for the same cost as Salad in the form of a '28core/32GB RAM/2x3080 Ti 12GB' on Vast.
So, what is your differentiation? Why would I ever use Salad?
There was a time I was running QDR Infiniband (40G) at home while everyone else was still dreaming of 10G at home because the adapters and switches were so expensive.
On the used market I'm sure you can get some decent kit for a homelab. But if you want a production environment, with anything close to the latest speeds (≥100G), you can't get gear for love or money: all the Big Players are taking all the inventory, so lead times are slightly ridiculous (assuming anyone has stock).
We were looking at IB in Feb 2022, and the earliest for IB switches was Nov-Dec 2022. We ended up going a different route for other reasons, so I haven't kept on lead times, but it would not be surprising if things were roughly the same.
Also hard to get from vendors: wait lists for >100G switches is a long time. No one has them in stock: probably because LLM is what all the cool kids are doing, so all inventory/production has been spoken for.
Any particular reason? I know latency is lower, which is good for MPI and such, as is RDMA.
But for general use, why should folks care about IB instead of Ethernet?
Things like "Datacentre Ethernet" and RDMA over Converged Ethernet (RoCE) are basically Infiniband anyway (you will find that such features are only available on Connect-X HBAs from Mellanox which can generally run in 100GE or IB mode).
The main reason these things matter is they are circuit switched and generally combined with network topologies that ensure full bisectional bandwidth (or something that is essentially good enough for the communication pattern the machine is designed for). This means the network becomes "lossless" once it's properly verified as online and circuits are setup. For storage this is essentially unmatched, it's equivalent to Fiber Channel but way faster and converged so you can also run your workload over the same links rather than having separate storage and network links.
If China invades Taiwan, GPUs will be the least of our concerns.
>> which appears increasingly likely
It's not likely. China gains from the sabre rattling and politicking. They've seen what happens when a bully country attacks a smaller country.
The leaders of the CCP have everything they could possibly want in this world. They won't risk it all for an uncertain outcome attacking Taiwan.
They don't have Taiwan. Historically, they think Taiwan is not a distinct country, but a splinter off of China that needs to be brought back into the fold again.
Just the same bullshit Putin thinks about Ukraine (and most probably also about Belarus). The only thing saving Taiwan from a Chinese invasion is just how bloody Putin's nose got in Ukraine - and that's without the defense pact Taiwan has with the US.
The official reasons given differ but generally are just an excuse to drive the nationalism which is the goal of the war in the first place. China's government seems to be doing well enough right now to not need that trump card.
While I agree that the Iraq war was completely out of any legal boundaries (unlike Afghanistan), it had never been the intention of the US to pull a land-grab war and turn Iraq into the 51st US state. The invasions of Putin however clearly were, and so will a potential Chinese invasion into Taiwan assuming they'll stick to their actions in Hongkong.
Clearly the incentives and motivations of people running oppressive regimes are slightly different.
China's military has been built with the explicit long term goal of invading Taiwan. They drill for it, with pretty provocative displays - e.g. flying ballistic missiles over the island last year.
I think they will continue to build their military until success is assured, winning the battle before it's fought, per Sun Tzu. That's the best way to fight: make resistance obviously pointless.
Russia, to take a different example, never had enough troops to occupy Ukraine. They had ~130k troops on the borders. Per [1], crushing the Czechoslovakia 1968 uprising and maintaining order afterwards required 500k troops for 14 million people, and that was without state military resistance. Ukraine had 3x that, Russia would have needed an army on the order of 1.5 million. So the plan was to fly in and take Kyiv in a decapitation attack, but it failed, and here we are.
China has an army of 2 million people actively serving. Taiwan is 23 million people. I think it's a matter of when, not if, as long as Xi is in the ascendant, and isn't too distracted by the growing problems of the lopsided debt-laden, export-focused, manufacturer subsidizing mercantilist Chinese economy.
[1] https://foreignpolicy.com/2022/02/03/russia-couldnt-occupy-u...
How many are assigned to area around Taiwan?
But also, the number is kinda pointless as they need to get to Taiwan somehow unless they want to swim across.
ChatGPT is cool, but not cool enough for me to be OK with the ridiculous prices Nvidia can get away with on GPUs nowadays.
Consumer GPU sales are at the lowest point they've been in decades, but Nvidia doesn't care. Their profits are higher than ever because everyone and their dog is paying out the nose to build up their AI nonsense.
If you do care about those things, then you can pay for them by buying a 3000 or 4000 series card. They're a bit more expensive.
I don't think Pascal cards will be relevant for AAA games much longer now that AAA games are starting to skip last-gen consoles. Ubisoft showed of their Avatar and Star Wars games this week, and both are shipping with ray-traced global illumination.
It'll incentive GPU development, and the more new/more-capable GPU models come out, the cheaper older/less-capable models will get.
We've seen this cycle again and again with all sorts of computer hardware, from hard drives to monitors, CPU's, to memory, to graphics cards themselves. With some notable exceptions, the power and capacity of most hardware consumers can afford today dwarfs what was affordable a decade or two ago.
The same will happen to GPUs if demand remains high. So bring it on!
https://cloud.google.com/compute/gpus-pricing
Am I misunderstanding the GCP pricing or how this analysis was done?
So the on demand price is not really relevant, maybe thats why they excluded it.
Instances from 1 to 8 x H100 GPUs, paid by the hour
That'll buy you about 13 RTX 4090's outright.
I think the price is quite high to be honest, although probably realistic. Wouldn't expect competitors to be able to offer something better and data centers need to calculate for maybe newer models with even better performance.
$1.89 / hour for 100% upfront
$2.04 / hour for 80% upfront
$2.15 / hour for monthly payment
Same I guess if they could use evaporation to do some of the cooling.
I just took a bunch of GPUs out of storage and racked up a machine yesterday. From 0 to Linux/Docker install to Stable Diffusion / llama in about 2 hours.
Facebook/Meta used 8,000 A100s to train LLaMA (for example).
If you’re doing anything with an LLM it’s fine-tuning and there’s a new approach nearly weekly to do it better, faster, cheaper, and easier on 24GB cards which can be had for less than $1000.
Same for inference.
I agree, and I have a 3090 for that purpose, and once wrote a tutorial for others wanting to do ML stuff on a GPU at home rather than rent from a cloud provider.
But a consumer GPU (or eBay older Tesla card) can't do everything that a rental pool of H100 and A100 can do, and I and other readers here will sometimes want to do those other things.
I didn't want to confuse people that all they'd need was to buy up a bunch of random retired Ethereum miner GPUs, no matter what they wanted to do with ML.
The other benefit I'd add with buying your own GPUs is availability. They are yours and always yours, in a commercial application with deadlines, etc it's a real risk to depend on being able to get the necessary on-demand GPU compute on various cloud platforms at any point in time. There is nothing worse than logging into a cloud provider console and seeing "no availability" when you really need to get something done. For me personally this is what pushed me to buying vs cloud because I ended up in scenarios where Vast.ai was the only option left and I haven't had the best experiences with Vast.ai in terms of reliability and performance (I'm pretty sure many of the benchmarks are gamed, although I'm not sure how).
Speaking of performance, I've also seen very real issues with virtualized CPUs, what I assume is network attached storage, etc feeding data to high end GPUs fast enough (again, noted elsewhere in this thread). In benchmarking that I've done with various cloud providers unless you go for the much more expensive options on GCP and elsewhere with directly attached NVMe storage a single NVMe drive and decent CPU in a workstation will run circles around many of these cloud providers.
In any case I think when you're talking A100s in the thousands my point remains. No one is just showing up cold to a cloud provider from a website link and spending at least tens of millions of dollars.
We are both running Stable Diffusion with Dreambooth etc., Llama (well, he is running Llama, I am running 4-bit Llama), Deep Floyd and so on. Both machines are running Ubuntu, don't know what distros are good for servers.
Incidentally, NVDA closed at 429.97 today, an all time high (it was 108.13 last year, after they were banned from selling A100s/H100s to China - it has almost quadrupled in price in a year).
We're a distributed GPU cloud with 10k+ consumer GPUs on our network. We can crank out 4500+ stable diffusion images per dollar (highest in the market), all on consumer 3060s.
The high-end GPUs definitely have a role but for a bulk of production/inference serving can be done on consumer GPUs at 80-90% of the cost.
Serious question. I think most companies without the top talent would do better using an API product.