AMD MI300 performance – Faster than H100, but how much?
semianalysis.com
semianalysis.com
They still need a usable GPGPU language to glue all that hardware together, too.
Like AMD announced today, and is shown in the article we're commenting on, AMD now is 1.3 times faster at FP8, INT8 flops with and without sparsity, meaning their "transformer engine" is 1.3 times faster.
The rest of your comment doesn't make sense. See the other comments on software in this thread.
That seems like a minor bump compared to the new H100, but sometimes the quality of your training, or the speed of your inference, is hugely effected by small thresholds of VRAM capacity per card. The MI300X can do things the H100 cannot, and judging by the sheer size of the silicon, do them reasonably quickly.
If they were going to do it on the consumer side, they would have done it already. And I would own a 7900 48GB instead of a 3090, and probably have debugged ROCm on several projects by now :/
Also, while this wasn't in the presentation, CUDA can be translated to HIP using automated tools. and - CuPBoPAMD is a CUDA translator that translates CUDA programs at NVVM IR level to HIP-compatible IR that can run on AMD GPUs (https://dl.acm.org/doi/pdf/10.1145/3624062.3624185)
AMD made a presentation on their AI software strategy at Microsoft Ignite two weeks ago. Worth a watch for the slides and live demo
Not just automated, but also open source.
They really couldn't come up with a better name than whatever that is?
Not sure when Microsoft will just have their own TPU accelerator but I wouldn’t be surprised if the announce it sometime.
https://news.microsoft.com/source/features/ai/in-house-chips...
I was not aware of any efforts to purchase or deploy MI300s at OCI.
Looking forward to trying out the MI300x.
If AMD sells say $3 billion USD into that market, is that a big net positive for them?
This is the same reason why they were so successful with EPYC. It's easier to find eight small 8-core chips with high clocks/low-power than to find one 56-core chip with that same high-clock/low-power set which is why Intel sold so many cut-down chips while AMD just sold their small amount of defective chips as 6 and 12-core consumer chips.
Nvidia will have a harder time shipping defective because everyone wants the best chips with the best performance/watt to save money and do things fast. Outside of that, there's not a huge market for defective 800mm2 chips and the massive GPU the come on. AMD can sell their defective units a couple chiplets at a time to non-AI customers in laptops and workstation cards.
We're not just talking defective, we are talking about a silicon lottery on performance and it varies between every single one.
Suddenly potentially having a lot less backlog might screw up Nvidia faster that it "helps" AMD too.
Theoretical numbers are nice, but they are just that, ethereal numbers on a piece of paper. And on a CUDA device, I know from practice that I can get at least 90% of the memory bandwidth.
https://dl.acm.org/doi/pdf/10.1145/3624062.3624203
Table 6
Is MI300 also a variant of their customer GPUs?
https://www.nvidia.com/en-us/data-center/nvlink/
I'm super curious to see how their networking results compare between transports when real testing results between large supercomputer-ish farms get out there.
Also, as AMD and HPE power some of the most performant supercomputers in the world, they wouldn't have won those contracts if their networking was subpar. Those use slingshot.
You may also be interested in reading about https://ultraethernet.org/ .
https://www.amd.com/en/products/software/rocm.html
Before that, they were pushing heavily for standard OpenCL, but that failed because the hardware wasn't as competitive and the ecosystem/tooling barren.
Yes, “just use libraries”. But as Andrej Karpathy used to say, “I don't need some library holding my hand and providing abstractions. Real men command GPUs with their own raw kernel code”.
(And, incidentally, there are much less libraries for ROCm, for this reason. Somebody should write them, and why do that with something that takes 5x more code to do the same thing?)
To get a picture of the current state which has changed a lot this MS Ignite presentation may be of interest. https://youtu.be/7jqZBTduhAQ?t=61
OK, I retract my statement.
And there's MI300A which has a combined CPU+GPU with shared memory space which already has supercomputer customers. You can say it's DOA relative to Hopper but MI300 will be AMD's most successful GPGPU yet.
If they can’t afford to out R&D Nvidia today, how could they afford to fund a competitor? In the past 20 years the only serious new entrant to the discrete GPU market has been Intel and they are also far behind Nvidia and CUDA.
I was actually motivated to contribute to the AMD drivers at some point and did land a couple of patches, but wasn't looking to switch career paths. The drivers have become quite good without me anyway, so no regrets :)
At the very least, pytorch has to work effortlessly across GPUs, down to the less popular and odd workarounds people build into their pyotrch code.
Even if RocM reaches parity, the first mover advantage for Cuda is too large. RocM porting has to be literally effortless. Everything else is DOA.
And the large companies (Meta, Microsoft) buy these GPUs by the thousands upon thousands.
Having a small team of engineers spend a couple months porting code over is well worth it for even modest cost reductions.
These may not be useful to smaller companies working in the AI space, but odds are to sellout, all AMD really needs are the sales contracts that they've already announced.
That being said, if you a cloud provider and need to scale up a bunch of basically similar transformers models, then it should be an easy sell for AMD.