H100 GPUs Set Standard for Gen AI in Debut MLPerf Benchmark
blogs.nvidia.com
blogs.nvidia.com
An equivalent number of V100's (GPT-3's original GPU) would've taken about 36 days [0].
0 – https://www.reddit.com/r/GPT3/comments/p1xf10/comment/h8h3sl...
That’s why these clusters are so valuable, even with Nvidias margins they are still cheaper than using less compute for longer.
that means they might pass a few batches through a GPT-3 sized random init network and time it
https://twitter.com/abhi_venigalla/status/167381386318645248...
(300BT / 1.2BT) * 11 min * (1 hr / 60 min) = 45.8 hr
Still pretty incredible. That's an 18.8x speedup over 36 days.An eight GPU DGX-1 server cost ~149k$ back then (googled news postings). A current gen DGX H100 is 520k$ with 5 years of support. Of course it holds 5x the memory, plus GPUs and interconnect are much faster. But when comparing costs, take price hikes into account.
For writing code you don't care about feeding world history to your model. So a smaller model might be better at a specialized task
Sure, having a big multi-modal-model is great, but by having specialized models you can spread tasks better
Just to compare the GPUs, TF32 Tensor processing went from ~125 TFlops to ~990. It then looks like they also dropped the precision to FP8, which gives you another 4x win.
What's interesting is to look at how we're progressing in performance over time. In some sense, a bit slow?
A V100 costs $10k at release; an H100 seems to be $40k.
So we've only managed to halve the cost of a flop in 5 years. That seems.. much slower than what Moore's Law would have suggested.
On a (minor) side note, it seems that $1.00 from 2018 is worth ~$1.20 in 2023. I wish more cost comparisons included inflation, because the past few years have had a lot of it.
You need to be an order of magnitude higher in power usage before it really starts mattering. e.g. a Tesla3 consumes around 15 KW (20x the H100) while driving.
That said it looks like flops/watt dropped by only 3.4x, which is also sub Moore's Law (3 years to halve power consumption)
You couldn't replicate the scale of compute this allows no matter how many V100's you had.
A huge amount of cost here is embodied in networking and memory.
If one were to design a chip that cared only for flop/$ without caring for all of the interconnect and memory, then the 4090 is a much fairer comparison, and even then that card isn't designed for a flop/$ optimisation.
Moore's law says nothing about price or performance.
H100 has about 4x as many transistors as V100, which is pretty close to what Moore's law would predict.
Moore's law is about doubling of transistor count for the same price. At least that's always been my understanding.
EDIT: I decided to look it up. Heres the original 1975 statement from him that led to the law:
"The complexity for minimum component costs has increased at a rate of roughly a factor of two per year. Certainly over the short term this rate can be expected to continue, if not to increase. Over the longer term, the rate of increase is a bit more uncertain, although there is no reason to believe it will not remain nearly constant for at least 10 years."
So yup, it's about transistor density vs price.
The V100 had up to 32gb ram, the H100 is 188gb. While there’s obviously transistors in ram, the counts being compared are for the GPU itself. I’d argue that a big chunk of the price difference in the V100 vs H100 is RAM.
But that's the same Moore's Law problem when using the 'price' definition.
6x RAM 5 years later = 4x the cost? That's even slower than the FLOPS gains.
Roughly the same tech advances apply to making all three kinds of chips so I think it applies (somewhat) equally to CPU, GPUs, RAM and even SSDs etc.
Bearing in mind that ram lags a few generations behind CPU and GPU (I think DDR5 is 12nm vs latest GPUs around 4nm). And also that Moore's law is not a physical law, even when it was in full swing it was only ever meant as a rough guideline for what to expect over a couple of years period.
If you bought the $2/hr H100 instance they offer, that would cost ~$330k. Pricey, but not too bad.
My bigger hope is that with cheaper compute we can see more architecture search and designs. A lot of different architectures are relatively unexplored due to computational constraints and are typically performed by smaller labs so they don't scale and it is kinda hard to compare models when we're just looking at performance benchmarks and not considering other factors. We definitely don't want big labs to railroad our research directions. Feels weird that a huge amount of NLP is based on using pretrained models and tuning them. Vision is going this way too. You basically can't get published without being SOTA so you basically have to modify an existing model or have a multi-million dollar lab and train from scratch. Really weird to expect academia to compete with big labs and really weird to not let academia take "bigger risks" and explore less popular areas. It is vital to our research path that we don't force everything onto a single track.
3,584 GPUs at $30,000 is $104,550,000 USD only in GPUs.
Or some other 'trust but verify' situation where the suggested action is still validated against business rules that have some notion of consumer protection.
On the side of hope, if somebody were to demonstrate how to use a bunch of GPUs to make the e-coli bacteria live for a little longer using an approach that has a semblance of generality, I guarantee that a lot of old people would be willing to part with their fortunes in exchange for a sliver of hope for them or future generations.
There is serious research going down right now, but you don’t see you typical medical research department at your university buying ads on Instagram.
FluidStack or Lambda Labs
If you need to rent a supercluster and you’re not tied to one of the big 3 clouds, then talk with FluidStack, Lambda, Oracle, maybe CoreWeave.
Tl;dr of the good ones: FluidStack and Lambda for H100s (1x instances), Runpod for A100s.
Data parallel models scaled up for training and then could run on individual chips, but these massive model parallel models require a couple of chips directly linked together even to do inference.
So the idea that a competitor could come in with a simple, cheap inference chip doesn't really work.
The problem remains the lack of software support.
But I don’t understand why software is a problem for them with their deep pockets. It can’t possibly be dearth of talent, or that it is expensive. Here in Europe a good software engineer earns half of what a mediocre software engineer earns in USA, to say nothing of India or China. They could just hire a bunch of teams and up their software game.
They also published MLPerf results today: https://habana.ai/blog/gaudi2-demonstrates-competitive-llm-p...
With 384 Gaudi2s they did the LLM task in 312 minutes, compared to 46 minutes for 768 H100s. It’ll come down to cost, but given the H100 is a process node or two ahead (and much more expensive I imagine?), Intel is actually much closer than I had realized. They’re the only other to submit an MLPerf result for the LLM task, I think. All credit of course to Habana Labs which was acquired by Intel.
But that said, Intel suffers from the same fundamental problem as AMD: They aren't Nvidia and they can't CUDA.
Probably better to aim for PyTorch compatibility. In practice, that's how most AI programmers interact with their GPUs.
Its just that current supply is extremely short, so the H100s end up only available to big buyers. But that will be resolved in time. Nvidia wants every university lab to have a H100 so no competitor sneaks in there.
How much do these GPUs cost?
H100s are around $30-33k at the IT hardware resellers CDW and SHI (https://www.cdw.com/product/nvidia-h100-gpu-computing-proces..., https://www.shi.com/product/45671009/NVIDIA-H100-GPU-computi...)
Supermicro’s HGX H100 8x GPU server is $297k at the reseller Dihuni (https://www.dihuni.com/product/supermicro-8125gs-tnhr-server...)
DGX H100 is $521k at the reseller Insight (https://www.insight.com/en_US/shop/product/DGXH-G640F+P2CMI6...)
The DGX GH200 might cost in the range of $10mm-20mm or more (A guesstimate ballpark from an exec at a cloud company I talked with)
If anyone wants to pre-review the post and can offer thoughtful comments, my email's in my profile.
And to clarify the difference between all of these product names, I put together this diagram - https://gpus.llm-utils.org/dgx-gh200-vs-gh200-vs-h100/. I don't have the HGX H100 or the DGX H100 on there though - but the HGX H100 is a reference platform for OEMs to design and make H100 based servers with either 4x H100s or 8x H100s (https://nvdam.widen.net/s/5kgbjq2v2t/hpc-hgx-h100-datasheet-...) and the DGX H100 is the official Nvidia server with 8x H100s (https://resources.nvidia.com/en-us-dgx-systems/ai-enterprise...).
Additionally, code for the actual submission is available here https://github.com/mlcommons/training_results_v3.0/tree/main...
[0] https://ourworldindata.org/grapher/carbon-intensity-electric...
It amazes me that computational feats of this magnitude can be so energy efficient in the scheme of things.
[0] https://www.epa.gov/greenvehicles/tailpipe-greenhouse-gas-em...
[1] https://www.epa.gov/greenvehicles/tailpipe-greenhouse-gas-em....
[0]: https://www.carbonindependent.org/22.html#:~:text=At%20a%20c...
Will you join the resistance?
LOL, I welcome our new overload.