Nvidia Hopper GPU Architecture and H100 Accelerator
anandtech.com
anandtech.com
> 80 billion transistors
> Hopper H100 .. generational leap
> 9x at-scale training performance over A100
> 30x LLM inference throughput
> Transformer Engine .. speed .. 6x without losing accuracy
So another monster chip - same size of the Apple M1-max thingy ..
I guess it comes down to pricing. The A100 is already ridiculously expensive at $10K. They can this one at $50K and it would sell out?
If you're interested in giving us a shot, feel free to shoot me an email at mike at crusoecloud dot com.
P100 -> V100 -> A100 -> H100
This is not confusing at all.
https://en.wikipedia.org/wiki/Category:Nvidia_microarchitect...
Fermi -> Kepler -> Maxwell -> Pascal -> Volta (HPC only) -> Turing -> Ampere -> Hopper (HPC only?) -> Lovelace?
Nvidia's accelerators are just called <Architecture Letter>100 every time so if you don't remember the order of the letters it's not obvious
They could have just named them P100, V200, A300 and H400 instead
However, is the number superfluous at this point?
J3710, 7th Gen J3060, 8th Gen
J4205, 8th Gen J4125, 9th Gen
i3-5005U, 5th Gen N5095, 10th Gen
i7-3770, 3rd Gen 3865U, 7th Gen N3060, 8th Gen
In general, complaining about naming is peak bikeshedding for the tech-aware crowd. There are multiple naming schemes, all of them are reasonable, and everyone hates some of them for completely legitimate reasons (but different for every person). And the resulting bikeshedding is exactly as you'd expect with that.
The underlying problem is that products have multiple dimensions of interest - you've got architecture, big vs small core, core count, TDP, clockrate/binning, cache configuration/CCD configuration, graphics configuration, etc. If you sort them by generation, then an older but higher-spec can beat a newer but lower-spec. If you sort by date then refreshes break the scheme. If you split things out into series (m7 vs i7) to express TDP then some people don't like that there's a bunch of different series. If you put them into the same naming scheme then some people don't like that a 5700U is slower than a 5700X. If you try to express all the variables in a single name, you end up with a name like "i7 1185G7" where it's incomprehensible if you don't understand what each of the parts of the name mean.
(as a power user, I personally think the Ice Lake/Tiger Lake naming is the best of the bunch, it expresses everything you need to know: architecture, core count, power, binning, graphics. But then big.LITTLE had to go and mess everything up! And other people still hated it because it was more complex.)
There are certain ones like AMD's 5000 series or the Intel 10th-gen (Comet Lake 10xxxU) that are just really ghastly because they're deliberately trying to mix-and-match to confuse the consumer (to sell older stuff as being new), but in general when people complain about "not understanding all those Lakes and Coves" it's usually just because they aren't interested in the brand/product and don't want to bother learning the names, and they will eagerly rattle off a list of painters or cities that AMD uses as their codenames.
Like, again, to reiterate here, I literally never have seen anyone raise AMD using painter names as being "opaque to the consumer" in the same way that people repeatedly get upset about lakes. And it's the exact same thing. It's people who know the AMD brand and don't know the Intel brand and think that's some kind of a problem with the branding, as opposed to a reflection of their own personal knowledge.
I fully expect that AMD will release 7000 series desktop processors this year or early next year, and exactly 0 people are going to think that a 7600 being newer than a 7702 is confusing in the way that we get all these aggrieved posts about Intel and NVIDIA. Yes, 7600 and 7702 are different product lines, and that's the exact same as your "but i7 3770 and N3060 are different!" example. It's simply not that confusing, it takes less time to learn than to make a single indignant post on social media about it.
Similarly, the NVIDIA practice of using inventors/compsci people is not particularly confusing either. Basically the same as AMD with the painters/cities.
It's just not that interesting, and it's not worth all the bikeshedding that gets devoted to it.
</soapbox>
Anyway, your example is all messed up though. J3710 and J3060 are both the same gen (Braswell), launched at the same time (Q1 2016), that example is entirely wrong. J4125 vs J4205 is an older but higher specced processor vs a newer but lower spec, it's a 8th gen Pentium vs a 9th gen Celeron, like a 3100X vs a 2700X (zomg 3100X is bigger number but actually slower!). And the J4125 and J4205 are refreshes of the same architecture with legitimately very similar performance classes. i3 and Atom or i7 and Atom are completely different product lines and the naming is not similar at all there, apart from both having 3s as their first number (not even first character, that is different too, just happen to share the first number somewhere in the name).
Again, like with the Tiger Lake 11xxGxx naming, the characters and positions in the name have meaning. You can come up with better examples than that even within the Intel lineup. Just literally picking 3770 and J3060 as being "similar" because they both have 3s in them.
The one I would legitimately agree on is that the Atom lineup is kind of a mess. Braswell, Apollo Lake, Gemini Lake, and Gemini Lake Refresh are all crammed into the "3000/4000" series space, and there is no "generational number" in that scheme either. Braswell is all 3000 series and Gemini Lake/Gemini Lake Refresh is all 4000 series but you've got Apollo Lake sitting in the middle with both 3000 and 4000 series chips. And a J3455 (Apollo Lake 1.5 GHz) is legitimately a better (or at least equal) processor to a J3710 (Braswell 1.6 GHz). Like 5700U vs 5800U, there are some legitimate architectural differences behind hidden behind an opaque number there (and on the Intel it's graphics - Gemini Lake/Gemini Lake Refresh have a much better video block).
(And that's the problem with "performance rating" approaches, even if a 3710 and a 3455 are similar in performance there's still other differences between them. Also, PR naming instantly turns into gamesmanship - what benchmark, what conditions, what TDP, what level of threading? Is an Intel 37000 the same as an AMD 37000?)
It is nice when products have a naming scheme where natural ordering of the name maps to performance.
We keep using code names in discussions because the actual names are ass backwards and not very descriptive.
Volta was HPC only while Turing was consumer line, I guess that may have something to do with the out of order naming sequence.
Then we had Ampere which was dual market. And then split again with Lovelace on the consumer side and Hopper for data centers.
Just guessing here, obvi.
Schedule me on an H100 and I promise I won't mind the "confusing" naming.
And for your typical dev - they'll interact with the GPU through a cloud provider, where they can easily know that a G5 instance is newer than a G4 one.
The degree to which energy and capital costs can be optimized will determine how large they can go.
We're at a point we we're turning into a computation driven society and computation is becoming a globally relevant power consumption aspect.
> global data centers likely consumed around 205 terawatt-hours (TWh) in 2018, or 1 percent of global electricity use
And that's just data centers, if you add all client devices you probably double that.
Plus that number will only continue to grow.
Why are people so dumb when it comes to planning for the future? Does it require a 1973 oil crisis to make people concerned about potential issues? Why can't people be preventative instead of reactive? Isn't the entire point of an engineer to optimize what they're building for the good of humanity?
Yeah, I strongly agree. While Nvidia is working on better hardware (and they're doing a great job at it!), we believe that better training methods should be a big source of efficiency. We've released a new PyTorch library for efficient training at http://github.com/mosaicml/composer.
Our combinations of methods can train CV models ~4x faster to the same accuracy on CV tasks, and ~2x faster to the same perplexity/GLUE score on NLP tasks!
The big welcome surprise for us the secure virtualization. Outside of some limited 24/7 ML teams, we mostly see bursty multi-tenant scenarios for achieving cost-effective utilization. MiG etc static physical partitioning was interesting -- I can imagine cloud providers giving that -- but more dynamic & logical isolation, with more of a focus on namespace isolation, is more relevant to what we see. Once we get into federated learning, and further disintermediations around that, even more cool. Imagine bursting on 0.1-100 GPUs every 30s-20min. Amazing times!
TF32 ....... 1,000 TFLOPS (tensor core)
FP64/FP32 ... 60 TFLOPS
I am more interested in the 144-core Grace CPU Superchip. nVidia is getting into the CPU business...That's still an amazing amount of power though. I can't help thinking about the kind of sear you could get on a steak with 40kW of power :D
If so, it may be a reasonable mix.
They'd need fab capacity first. I wouldn't count on it any time soon, and chip gens have short lives.
I share your suspicion that fp64 and ML workloads are distinct but can see each running on the same cluster at different times.
Looking good.
When measured at smaller #s of GPUs which are more realistic, the speedup is somewhere between 3.5x - 6x. See the GTC Keynote video at 38:50: https://youtu.be/39ubNuxnrK8?t=2330
Based on hardware specs alone, I think that training transformers with FP8 on H100 systems vs. FP16 on A100 systems should only be 3-4x faster. Definitely looking forward to external benchmarks over the coming months...
If 1000 TFLOPS is possible to do in inference time then im speechless
But keep in mind the model won't fit on a single H100 (80GB) because it's 175B params, and ~90GB even with sparse FP8 model weights, and then more needed for live activation memory. So you'll still want atleast 2+ H100s to run inference, and more realistically you would rent a 8xH100 cloud instance.
But yeah the latency will be insanely fast given how massive these models are!
Sounds doable in a generation or two.
1) NVIDIA will likely release a variant of H100 with 2x memory, so we may not even have to wait a generation. They did this for V100-16GB/32GB and A100-40GB/80GB.
2) In a generation or two, the SOTA model architecture will change, so it will be hard to predict the memory reqs... even today, for a fixed train+inference budget, it is much better to train Mixture-Of-Experts (MoE) models, and even NVIDIA advertises MoE models on their H100 page.
MoEs are more efficient in compute, but occupy a lot more memory at runtime. To run an MoE with GPT3-like quality, you probably need to occupy a full 8xH100 box, or even several boxes. So your min-inference-hardware has gone up, but your efficiency will be much better (much higher queries/sec than GPT3 on the same system).
So it's complicated!
I really do wonder how much more you could squeeze out of a full pod of gen2-H100's, obviously the model size would be ludicrous, but how far are we into the realm of dimishing returns.
Your point about MoE architectures certainly sounds like the more _useful_ deployment, but the research seems to be pushing towards ludicrously large models.
You seem to know a fair amount about the field, is there anything you'd suggest if I wanted to read more into the subject?
A pod of gen2-H100s might have 256 GPUs with 40 TB of total memory, and could easily run a 10T param model. So I think we are far from diminishing returns on the hardware side :) The model quality also continues to get better at scale.
Re. reading material, I would take a look at DeepSpeed’s blog posts (not affiliated btw). That team is super super good at hardware+software optimization for ML. See their post on MoE models here: https://www.microsoft.com/en-us/research/blog/deepspeed-adva...
Are these techniques for specific architectures or can they be made generic ?
See the paper here, Figure A28: https://kstatic.googleusercontent.com/files/b068c6c0e64d6f93...
But if your downstream task is simple, like sequence classification, then it may be possible to compress the model without losing much quality.
> For now, DPX ISA details are available to early access partners. We anticipate broader info availability aligned with CUDA 12.0 release later this year.
https://www.wolframalpha.com/input?i=30%5E3+cubic+nanometers...
Edit: big TDP, though.
Edit; many of these speeds are low precision which is less useful outside of deep learning, but the higher precision matmul ops in the tensor cores are still very fast and very useful for wide variety of tasks.
The FP64 matrix-multiplication is only 60 TFlops, no where near the advertized 1000 TFlops. TF32 matrix-multiplication is a poorly named 16-bit operation.
I'm on Turing architecture so I've never used TF32. I've only used FP32 and FP16 but FP32 isn't supported by these tensor cores.
Given that it's 32-bit in memory (so all your data structures are 32-bit) and also that in my experience using it is very transparent (I haven't run into any numerical issues compared to full FP32), I think calling it a 32-bit format is a reasonable compromise.
Addition is done in 10-bit mantissa. So maybe TF19 might be the better name, since its a 19-bit format (slightly more than 16-bit BFloats).
Really, its a BFloat with a 10-bit mantissa instead of a 7-bit mantissa. 10-bit mantissa matches FP16, while the 8-bit exponent matches FP32.
So TF19 probably would have been the best name, but NVidia like marketing so they call it TF32 instead.
Yes, the system will read/write the 32-bit value to RAM. But if there's only 10-bits of mantissa in the circuits, you're only going to get 10-bits of precision (best case). The 10-bit mantissa makes sense because these systems have FP16 circuits (1 + 5-bit exponent + 10-bit mantissa) and BFloat16 circuits (1 sign + 8-bit exponent + 7-bit mantissa). So the 8-bit exponent circuit + 10-bit mantissa circuit exists physically on those NVidia cores.
-------
But the 'Tensor Cores' do not support 32-bit (aka: 23-bit mantissa) or higher.
Only the ppas from graphics-drivers work properly
My experience on windows is much more automatic and it never breaks anything. But I'd rather pay the price (installing on Linux) to avoid windows at all costs
I highly recommend sticking with one technique or the other; never intermix them.
I wish people would stop talking rubbish about NVIDIA's Linux support.
The effective mantissa is like FP16 but it's padded out to be the same size as FP32.
In other words, there's 1 sign bit, 8 exponent bits, 10 mantissa bits that are USED, and 13 mantissa bits that are IGNORED.
1 + 8 + 10 + 13 = 32
The 13 ignored mantissa bits are part of the memory image: they pad the number out to 32-bit alignment.
Practically speaking you have the right to do anything unless someone complains about it. A lot of popular figures, even those long dead, have estates and organizations that manage their likeliness and other related copyright and IP. IDK what the situation is in this case, but Nvidia may very well have paid for the name.
I don’t know anything about the state of Dojo, but Tesla was very hand wavy about their software stack during their presentation. And running AI algorithms efficiently on a piece of hardware is one of those things that many HW vendors have a hard time getting right.
NVIDIA DGX H100 (8x H100): 480 TF FP32 / 8 PF+ TF16 / 16 PF INT8 / 640GB HBM3 / 10kW
Dojo off-chip BW: 16 TB/s / 36TB/s off-tile
H100 off-chip BW: 3.9TB/s / 400GB/s off-DGX
They can make an AI factory that will fit many times the current internet throughput in a small room.
I'm super excited about what kinds of applications this could mean.
Obligatory Linus Torvals on NVIDIA: https://www.youtube.com/watch?v=_36yNWw_07g