Nvidia’s H100: Funny L2, and Tons of Bandwidth
chipsandcheese.com
chipsandcheese.com
For instance, onnx & PyTorch just died randomly and a few thousand dollars later we figured it was directly caused by incorrect assumptions made in the underlying implementations where the developers almost definitely did not have access to an H100 yet.
Dedicated accelerators do that already, e.g. google's TPUs, tesla's D1 or apple's neural engine. You must load the data into compute-unit local memory first before executing matmuls. Keeping the weights there and only piping the dynamic data through it saves memory bandwidth.
Something else that might be theoretically possible is -
Large array of FPGAs are apparently used to simulate and verify chips [1], can the same be done to run LLMs? Can we have 0.25 to 1 token per cycle, would the engineering effort be worth it, and would it be financially feasible from a TCO standpoint?
[1] https://www.servethehome.com/amd-vp1902-is-leviathan-fpga-do...
https://www.microsoft.com/en-us/research/project/project-bra...
edit: typo
Not to mention that GPUs already execute in-order (at least any that I’m familiar with). They do have multiple execution pipelines, but instruction fetch/decode is in-order unlike something like a typical modern high performance CPU.
I'm unsure if this would be much of a win.
What a bummer. But is this really true? I know Nvidia does not report the cl_khr_fp_16 extension, but I saw somewhere that you can still use fp16 types in your code. Has anyone tested this?
What does "enabled" mean in this context?
> We’re testing H100’s PCIe version on Lambda Cloud, which enables 114 of those SMs, 50 MB of L2 cache, and 10 HBM2 memory controllers. The card can draw up to 350 W.
> Nvidia also offers a SXM form factor H100, which can draw up to 700W and has 132 SMs enabled.
So i wonder if the number of enabled elements is due to a power supply or cooling constraint.
It's also possible the yields are not so great and then you have a limited number of good SMs per chip
H100 SXM5 has 132/144 SMs enabled. Also higher clocks, much higher TDP.
According to Tom's Hardware, one could find 80GB boards for around 3500$.
BTW: that number was obtained from converting Yen to Dollars, and other figures mention instead "over 30,000$".
:P
Sure, if you want 10x the memory bandwidth of a smaller card, that should be expensive.
But 80GB of GDDR6 would currently cost something like... $300. Or if you looked 1-3 years ago it would have been $1000.
GDDR6 is already designed so that you can attach 8 data lines per chip. A high end GPU with a 384-bit memory bus could attach 48 chips that way and have 96GB. Or exactly 80GB on a 320-bit bus.
$1000-1500 retail baseline for the GPU, $250-800 for extra RAM, $1000+ for the extra design hassle... I think you'd be able to buy that for $3500 if we had better competition.
For a lot of use cases, you want a balance between compute and memory. For big AI models outside of a datacenter, memory is far more important. It's worth putting half the budget into RAM chips if that means your model can fit, even if you "only" get 100 teraflops at FP16.
https://www.tomshardware.com/news/nvidia-hopper-h100-80gb-pr...
just five weeks old.
There are many other articles on Tom's Hardware, mostly old. The newest one (two weeks ago) is just a small piece that reveals that
> can barely render graphics [as in, real time 3D rendering] as they do not have enough special-purpose hardware [...] GH100 only has 24 raster operating (ROPs) units and does not have display engines or display outputs [...] One H100 board scores 2681 points in 3DMark Time Spy, which is even slower than performance of AMD's integrated Radeon 680M, which scores 2710