Nvidia H200 Tensor Core GPU
nvidia.com
nvidia.com
https://www.anandtech.com/show/21136/nvidia-at-sc23-h200-acc...
This is an H100 141GB, not new silicon like the Nvidia page might lead one to believe.
https://www.anandtech.com/show/18780/nvidia-announces-h100-n...
The HBM3e loadout is slightly different than H100 NVL's was going to be, but this definitely seems like a higher bin H100. It's basically as-if AMD had shipped a 7900 XT, then latter started selling the 7900 XTX; same chip, but they brought up all the memory controllers on this one.
Sometimes things really are compute bound, and sometimes you get a "big" workload that still fits nicely in the GPU's L2. Generative AI is mostly at the far end of "memory bound."
Some ML startups (like Graphcore) seemed to bet on large caches, sparsity and clever preprocessing instead of raw memory bandwidth, but I think their strategy was compromised when model sizes exploded. Even Cerebras was kinda caught off guard when their 40GB pizza was suddenly kind of cramped.
Memory-bound operations seem to rather consistently be the limiting factor in my personal ML research work, at least. It can be rather annoying!
High Bandwidth Memory > HBM3E: https://en.wikipedia.org/wiki/High_Bandwidth_Memory#HBM3E
The memory makers bump up the speed the memory itself is capable of through manufacturing improvements. And I guess the H100 memory controller has some room to accept the faster memory.
Is the error rate due to quantum tunneling at so many nanometers still a fundamental limit to transistor density and thus also (G)DDR and HBM performance per unit area, volume, and charge?
https://news.ycombinator.com/item?id=38056088 ; a new QC and maybe in-RAM computing architecture like HBM-PM: maybe glass on quantum dots in synthetic DNA, and then still wave function storage and transmission; scale the quantum interconnect
Is melamine too slow for >= HBM RAM?
> Optical tweezers: https://en.wikipedia.org/wiki/Optical_tweezers
> "'Impossible' photonic breakthrough: scientist manipulate light at subwavelength scale" https://thedebrief.org/impossible-photonic-breakthrough-scie... :
>> But now, the researchers from Southampton, together with scientists from the universities of Dortmund and Regensburg in Germany, have successfully demonstrated that a beam of light can not only be confined to a spot that is 50 times smaller than its own wavelength but also “in a first of its kind” the spot can be moved by minuscule amounts at the point where the light is confined
FWIU, quantum tunneling is regarded as error to be eliminated in digital computers; but may be a sufficient quantum computing component: cause electron-electron wave function interaction and measure. But there is zero or 1 readout in adjacent RAM transistors. Lol "Rowhammer for qubits"
Of course even If 300GB GPUs were available tomorrow, and you sold a million house to buy as many as that would allow it’d still take years to train once.
Maybe in some cases, but that doesn't even really matter since hardware support is poor.
> Our goal is to simplify and accelerate ML development by creating more interoperability between various ML frameworks (such as TensorFlow, JAX and PyTorch) and ML compilers (such as XLA and IREE).
From there, their goal would most likely be to work with XLA/OpenXLA teams on XLA[3] and IREE[2] to make RoCM a better backend.
[1] https://github.com/openxla/stablehlo
Other backends are also available, such as CPU-only training. And you can export networks in reasonably-standard formats.
nvidia's moat is much more mature framework support than AMD's cards; widespread popularity due to that good framework support, ensuring everyone develops on nvidia, thus maintaining their support lead; much faster performance than CPU-only training; and a price that, though high, is a lot less than an ML developer's salary.
If you need 24GB of vram and nvidia offers that for $1600 while AMD offers it for $1300, how many compatibility problems do you want to deal with to save a single day's wages?
But nvidia's moat is far from guaranteed. Huge users like OpenAI and Facebook might find improving AMD support pays for itself.
At that scale they may actually develop their own hardware a la Google TPU.
If you want to just focus on the AI problem and not on infrastructure, just use NVidia. If you want control and efficiency, design your own. AMD kind of falls in a weird middle ground with respect to the massive companies.
> Our goal is to simplify and accelerate ML development by creating more interoperability between various ML frameworks (such as TensorFlow, JAX and PyTorch) and ML compilers (such as XLA and IREE).
From there, their goal would most likely be to work with XLA/OpenXLA teams on XLA[3] and IREE[2] to make RoCM a better backend.
[1] https://github.com/openxla/stablehlo
Not that anyone cares, and everyone keeps using CUDA while simultaneously complaining about Nvidia GPU prices, as if those two things have nothing to do with each other...
(There are some annoying differences in the low-level implementations of OpenCL vs. Vulkan Compute, due to their being based on SPIR-V compute "kernels" vs. "shaders" respectively, that make it hard for them to interop cleanly. So that's why the choice can be significant.)
I did a bit of work in OpenCL almost 10 years ago, and found it decently portable on a range of NVIDIA GPUs as well as Intel iGPUs. On the high end I used something like the Titan X while on the low end it was typical GPUs found in business class laptops.
But my limited exposure to AMD was terrible by comparison. Even though I am away from that work now, I still tend to try to run "clpeak" and one of my simpler image processing scripts on each new system. And while I liked a Ryzen laptop for general use or even games, it seemed like OpenCL was useless there. It seemed my best option was to ignore the GPU and use Intel's x86_64 SIMD OpenCL runtime.
Also my fractal software incl OpenCL multi-GPU / mixed plaftorm rendering: https://chaoticafractals.com/
Both work on [ Nvidia, AMD, Intel, Apple ] x [ CPU, GPU ].
Some of the shared code here: https://github.com/glaretechnologies/glare-core
Don't let anyone tell you OpenCL is dead! Keep writing OpenCL software!!
Only C, C++ and Fortran were never taken seriously enough, other language stacks never considered.
Thus everyone that enjoyed programming in anything not C, with great libraries and graphical debuggers flocked to CUDA, now remains to be seen if SYCL and SPIRV will ever matter enough to regain some of those folks back.
There's relatively few people capable of implementing these frameworks without a solid cuda-like foundation, and those that do exist would need a very strong incentive to do it.
AMD is trying to catch up too, so far Nvidia still remains to be the leader, a few years ahead.
So if TPU clusters are priced right...
Nothing is insurmountable. :)
https://en.wikipedia.org/wiki/The_Innovator's_Dilemma
>It describes how large incumbent companies lose market share by listening to their customers and providing what appears to be the highest-value products, but new companies that serve low-value customers with poorly developed technology can improve that technology incrementally until it is good enough to quickly take market share from established business.
Given how large the prize is, the next chapter of chip development is likely to be nvidia vs state sponsored projects. China, in particular, will funnel further resources into acquiring this technology by any means necessary, including (more) industrial sabotage and outright theft. It's going to be interesting to see how this will play out. Up until a few years ago China was viewed as being a formidable competitor for projects of this nature, but as the country has moved to become increasingly authoritarian, so too have its decision making and execution declined in quality.
Diversifying manufacturing away from Taiwan makes their position more precarious, not less.
I never thought I'd root for Intel.
Intel is on track with their node rollout roadmap, according to their CEO - [1]
[0] - https://www.tomshardware.com/news/nvidia-ceo-intel-test-chip...
There aren't many many Gaudi/Instinct cloud offerings even though the market is accelerator starved.
Epyc took the performance crown from Intel. Games consoles have been AMD for ages.
AMD are competing with Intel and Nvidia simultaneously with fewer resources than either, having come back from near bankruptcy in recent memory.
There's been plenty of effort and execution from team red.
It's commercially unfortunate that the crypto and now deep learning crowd don't particularly value the flexibility or control that comes from an open source toolchain. Regardless, I don't think the Cuda moat will hold out.
Quite the contrary, they've turned around the company to focus on AI.
Legacy software projects are on hold and software developers moved to work on AI under a new VP (former Xilinx exec). They have purchased some startups to get experienced AI developers.
Here is Andrew Ng giving a positive evaluation of AMD's software efforts https://youtu.be/KDBq0GqKpqA?t=2359
And the B100 is farther away. Nvidia always doubles the memory of their cards like this mid generation.
Most of the GPU's are backordered bigtime but our vendors are chomping at the bit to sell us these.
If you want a consumer GPU, you can go for the RTX 4090 (24GB VRAM) or the A6000 Ada (48GB VRAM) if you are building a workstation.
If you really need to "experiment" on an A/H100, then you can rent it by the hour through a cloud provider like Runpod.
It still has a media encode/decode blocks. A big one, in fact.
As for the number, the die name counts down to 100 (with GA107, for instance, being a small GPU die and GA100 being the big one), and the big datacenter GPU as a product inherits the 100.
Also, it occurred to me that Nvidia does sometimes increment the die to 200 (EG GM200, as the Maxwell 100 series was a single small oddball die). Its possible that they "refreshed" the GH100 die and are codenaming it GH200.
https://en.m.wikipedia.org/wiki/Fahrenheit_(microarchitectur... -> https://en.m.wikipedia.org/wiki/Celsius_(microarchitecture) -> https://en.m.wikipedia.org/wiki/Kelvin_(microarchitecture) -> https://en.m.wikipedia.org/wiki/Rankine_(microarchitecture)
Becoming a successful cloud provider is far from trivial, even if you can offer technology no-one else has.
What they're doing is instead trying to make sure that their GPUs continue to be seen as the best option in the short/medium term (by having them accessible everywhere), and trying to commoditize their complement by giving small cloud providers disproportionate GPU allocations, which they hope will drive customers from the big providers to the smaller ones that a) aren't trying to build their own ML hardware, b) will have less negotiating leverage with Nvidia in the long term.