Nvidia Hopper Sweeps AI Inference Benchmarks in MLPerf Debut
blogs.nvidia.com
blogs.nvidia.com
You don't get these types of generational improvements in the CPU world these days. In some ways though,it’s better that the gains go to the GPU because, after years of talk, we actually do seem to be in the middle of some really awesome ML developments.
- Significantly-improved diffusion models (DALL-E 2, Midjourney, Stable Diffusion, etc)
- Diffusion models for video (see https://video-diffusion.github.io/, this paper is from April but I expect to see a lot more published research in this area soon)
- OpenAI Minecraft w/VPT (first model with non-zero success rate at mining diamonds in <20min)
- AlphaCode (from February, reasonably high success rate on solving competitive programming problems)
- Improved realism and scale for NeRFs (see https://dellaert.github.io/NeRF22/ for some cool examples from this year’s CVPR)
- Better sample efficiency for RL models (see https://arxiv.org/abs/2208.07860 for a recent real-world example)
I'd add GPT3 & Github Copilot, which my team and I use professionally. It's far from perfect, but it's a great GSD tool especially for stuff like Regex, bash scripts, and weird APIs.
reasonably high was 50% on 10 attempts, meaning success rate on first attempt can be as low as 5%, out of which who knows how many were leaked to training data.
Not fast enough (not at the price point) but getting there: of course since it's an entirely parallel problem, the real metric is cost-per-model since in reality you'd just buy time from the cloud.
I figure we're going to start seeing some big changes once cost-per-model puts it within reach of the hobbyist. If I today could buy a custom model for say, $1000 - that's accessible to the hobbyist experimentalist.
that has been disclosed :/
I think you have to compare it to how much ETH you could have mined as the opportunity cost of training a model vs other uses of a GPU
I was more contextualizing this, where a potential 75% increase in power (on a smaller node) to deliver 50% more performance is less impressive:
> I think it's even more impressive that the smallest increase vs. A100 is still 50% (for ResNet)
If I were deploying 8-GPU nodes today, I'm not sure whether I would want to pick between 4-GPU nodes or doubling my node TDP. The DGX A100 and H100 both have 8 cards, and 640GB of VRAM, but peak power increased from 6.2kW -> 10.2kW. If you're training giant models and constrained on VRAM, you'll potentially need just as many nodes as before just to fit your models into memory. Your training will be faster but it's hard to avoid the power increase.
The power increase is much more reasonable if you stick with PCIe SKUs (300W -> 350W, literally half the power draw of the SXM SKU), but you pay for that with 20% less compute and 33% less memory bandwidth.
That's misleading, as the GPU power utilization during ResNet is very likely not 100% for either the A100 or H100.
Remember the lower-TDP PCIe H100 has 20% slower compute and 33% slower memory than the SXM model, suggesting the increased power delivery from the PCIe (350W) to SXM (700W) model is a major factor in the performance even for H100 vs H100.
I don't think it's misleading to say that power is extremely likely to be a factor in the demonstrated performance increase from A100 to H100 for non-transformer workloads until proven otherwise. I don't think anyone here has a DGX H100 in hand yet to test this.
Edit: wait, also I looked closer at the numbers - in MLPerf 2.1, NVIDIA only submitted results for 1x H100. There's no 1x A100 result submitted for resnet50, just an 8x A100 number, which NVIDIA seems to have divided by 8 to get a "per accelerator" number to compare to their 1x H100. That doesn't feel like a clean comparison, as you can have performance loss when scaling. I'd rather see 1v1 or 8v8.
It might be easier to grasp if you use 1x. E.g. "it's 1x as fast" means it's exactly the same. You're applying the multiplier directly to get the new speed. 0.5x as fast means it's half as fast.
"It's 1x faster" means you add the result of the multiplication to the initial value. So it's twice as fast. 0.5x faster still means it's faster, you add 50% of the initial speed. This way you cannot really express that something got slower, except by using negative values, which might be rather confusing.
I think this might also work in most other languages originating in Europe.
To my ears, "3 times faster" has the same meaning as "3 times as fast".
Some people want to use language that is logical and consistent. For them, "3.5x faster" means the same as "350% faster" and "4.5x as fast". Others believe in convenience and redundancy. They think that "4.5x faster" means the same "4.5x as fast", because the numbers are the same. Many of them don't like expressions such as "X% faster" for X >= 100, because the numbers tend to be misleading.
I used to be in the former camp when I was young. Today I'm middle-aged and lazy. When I see "3.5x" in the text, I assume that it means "3.5x". I don't want to read the text carefully to determine that you actually meant "4.5x" when you wrote "3.5x". I interpret the "faster" part as redundancy. It tells me that we are talking about speed and that the thing we are talking about is faster than the baseline.
You could say we are mid-journey…
T4 Tesla gets 21,691 queries/second for for ResNet, compared to 81,292 q/s for the new H100, 41,893 q/s for the A100 and 6164 q/s for the new Jetson.
So you can expect maybe 15,000 q/s on a M1 Max. But some tests seem to indicate a lot less[2] - not sure what is happening there.
[1] Setup like this: https://github.com/nlothian/m1_huggingface_diffusers_demo
[2] https://tlkh.dev/benchmarking-the-apple-m1-max#heading-resne...
Source: https://www.reuters.com/technology/nvidia-says-us-has-impose...
As far as dedicated GPUs. There is so much additional work in the software side/driver side of thing unless there was a big concerted effort by a single company to build a solid foundation I see it as unlikely. Look at Intel struggling to get their GPUs out there.
Dedicated accelerated cards just for ML though? I believe some already exist/more will come.
If previously building a separate CUDA-like framework competed against paying a bit more to Nvidia, now this decision is infinitely simpler, since using Nvidia products is just out of the question.
Quite a gift to the Chinese if you ask me but US policy makers seem blind to this.
ARM Ltd is possibly one exception. Though after many many years of attempting high performance designs and coming out with mediocre designs they promptly had their doors absolutely blown off by Apple i.e., the US PA Semi design team, in its first outing (or second, if you count PA6T, which I guess you should, but the PA6T itself was superior in performance to ARM cores when it was released too, they just didn't play quite so obviously in the same markets). Japan does some boutique supercomputer CPUs I suppose although I don't think they're actually competitive so much as a protected and subsidized venture.
I don't know why. China, Russia, Japan, Europe have been attempting this for a long time with comparatively little success, then you get a startup in the US come out with something great.
It's interesting, it almost seems there are dynasties of some secret sauce, the recipe of which was discovered in 1960s and they only pass it down by word of mouth. There are so many of these design teams in startups and established companies you hear about with roots in DEC or Intel or IBM.
Maybe that's changing. Maybe it's less true with GPUs than CPUs. Not sure though, USA is certainly the center of the universe for all that stuff at the moment.
I don't know about this chip. Possibly this time it'll be different, unlike all the previous times it would be different but wasn't. I don't think that's been established yet though. And certainly it's not for CPU cores or even more general purpose GPUs.
It's already here. What the fuck are you talking about, "I don't think that's been established yet". That would have made sense in 2019, not in 2022.
Hardware side I think this is less true, but in general hardware folks and electrical engineers in general don’t get the absolutely outrageous SV salaries in SV.
Chinese Startup Biren Details BR100 GPU: https://www.hpcwire.com/2022/08/22/chinese-startup-biren-det...
So the company that has designed it will survive and it will grow and be able to design a next generation of GPUs, being protected from stronger competitors.
Had China decreed an overtax on imported datacenter GPUs to help its domestic producer, USA would have protested against such a protectionist measure. But now USA has done itself what was needed to protect the Chinese datacenter GPUs.