That being said, it’s a Grayskull Processing Unit. ;)
The H100s feature a minuscule amount of ROPs and TMUs (mainly for graphics); and doesn’t come with NVENC/DEC, etc. DirectX, Vulkan is not supported.
You’ll just see more and more tensor cores, optimised CUDA cores for bfloat, etc.
https://www.servethehome.com/wp-content/uploads/2023/10/NVID...
Tbh it’s mildly surprising they even removed NVENC considering the overall size of the chip (in the absolute we are only talking about low-single-digit mm2 savings) and then H100 is still advertised and has features targeting VM graphics/visualization still… remember they also still put a full graphics pipeline with ROPs/TMUs on the chip, just no actual display hardware.
These days, simulation runs basically nonstop on all the FPGAs the company is able to throw at it until tape out. Then, in post-silicon validation, you have engineering teams working 24/7 to validate the device on the bench. Finally, test engineering tries to achieve yield in HVM.
Anything that you don't need on the spec sheet or the designer insists on is cut to save time to market.
Apple initially developed OpenCL with Khronos as a similar-ish analog for GPU acceleration a while ago. There were a few partners that invested in it, but it suffered the same lack of demand as CUDA and languished for a while. Now Apple doesn't support OpenCL, so the purpose of a multi-platform acceleration library has kinda been scuttled. Nvidia played their cards well, and their competitors are going to feel the pain for a while unless they work together again.
Not much. GPUs became array processing engines in the early 2000s and the amount of residual fixed function graphics hardware is tiny now.
The biggest thing is that a lot of network evaluation is being done at lower precision (BF16/ FP8 / ternary) now, but my understanding is that training the networks still requires FP32. As these companies are really targeting the training market the high precision modes needed for graphics can't be thrown away. If a company was to target the end user, then you could optimise for lower precision, but end users find dedicated neural net acceleration a luxury they can get for free by buying a GPU.
I always assumed that was what killed Xeon Phi. Hard to justify a GEMM card in and of itself. However clever you get, NVIDIA will be back next year with twice as much bandwidth.
If you haven't yet, buy some bitcoin and stop letting the government steal your hard earned money through inflation.
1. The hardware isn't nVidia's primary moat. CUDA is. Currently all machine learning tooling assumes you have a CUDA device.
2. The hype cycle may be real, but nobody really knows how much better AI will become. The uncertainty is driving the hype (which is obviously on the optimistic side), but it could well be that models become so good that everyone "needs" an AI chip like how we "need" a smartphone today.
3. The gaming/graphics market is probably going to slowly shrink. There's only so many pixels you can push to the eye. There's very real diminishing returns there. At some point people literally won't be able to tell the difference between the newer/better/faster graphics and the older generation, at which point the GPU vendor's margins drop significantly.
People really don't appreciate this. So many times I see people comment that AMD should implement CUDA and not try to get software to move standard.
AMD have a massive task ahead because they NEED to get CUDA unseated just to play on a level playing ground, let alone win.
Melting connectors, obscene pricing, marginal performance increase, zero availability; then when hardly any of that was addressed, Ti versions release which are equally expensive, and equally lacklustre. Combine that with patchy support for frame generation and 30xx series cards still going really strong, and you’ve got not a lot of incentive to upgrade.
In units sold?
I’m sure they don’t mind having both revenue streams of course! But on the other hand, they have to serve two sets of interests.
I put those in quotes, because theoretically both can do training, it's just that grayskull isn't well suited for it, because of the internal fp19 format, and no good support for scaling to multiple cards. wormhole will have internal fp32, and is designed to scale out with across multiple cards, and servers.
At least that's how I understood it.
In LLM-land we've rapidly gone from training bespoke models to doing fine tuning to RLHF to zero-shot prompting. The better the underlying model the more you can do without additional training, so hardware that fails to scale up to the largest training runs will have limited practical utility despite technically supporting training.
And yes, I know, I've been working on LLM pretraining for about 4 years now, since 2020. The number formats themselves so far are mostly scale invariant or improve with larger scale - you can quantize a larger model and see less performance drop than a smaller model.
Indeed.
GPGPU means "using a GPU for things that are General Purpose (GPGPU)". The aforementioned IC is not a GPU that is used for general-purpose computations, but a specialized chip for AI computations.