AMD MI300X 30% higher performance than Nvidia H100, even with optimized stack
tomshardware.com
tomshardware.com
If you are a LLM startup blocked on compute, send me a dm
Tensorflow had first mover advantage but they also had first mover disadvantage. TF isn't that easy to write nor to debug. PyTorch got to see everything that was wrong with TF and fix it. No doubt Meta has put in lots of amazing effort. It is an incredible piece of software.
I'll also add that PyTorch does a lot of great stuff that people don't use or recognize. There are a lot of good distributional works out there for prob and stats people that give you cuda acceleration. Really a lot of numpy can be replaced with pytorch and its great. But be warned of FMA and other optimization differences.
Not true. PyTorch has had an MPS backend[1] for Apple Silicon for quite some time now. And lately, they've also added a HIP backend[2] for AMD hardware. I've written CUDA kernels for PyTorch in the past, it'd have been a lot easier if PyTorch were a thin wrapper over CUDA, but it isn't.
==> You can also say it's a lot of other things like ROCm. <==
I mean my intent was just to communicate that PyTorch runs CUDA in the backend, which was what I meant by API, because it seemed the person I was responding to wasn't aware of this backend code you're interfacing with. I thought adding the same phrasing about it also "being" ROCm and using "such as" and "other things," that others would have allowed anyone with more knowledge to infer that I'm aware of the generalization but I guess I was deeply mistaken given the replies that are certain that I can't read the pytorch homepage. Though I'm not quite sure how you came to the conclusion that I wasn't aware of ROCm support since I did explicitly mention it.
Do you have feedback for how I can better communicate? I seem to be running into failures like this a lot lately where I feel like I can point to where I clearly stated something but that doesn't matter because communication is about getting the person on the other side to receive the message. I've been scratching my head about this and could definitely use the outside perspective.
Obviously you can use cuda to write more than ML algos, but that’s its overwhelming use case.
Thus my point is if AMD ported PyTorch to their hardware (and PT is not cuda-only) it would create a more level playing field on which they could compete (in a cousin vs cousin battle) by eliminating the switching cost for the majority of developers.
This is the same reason cpu mfrs have compiler ports but generally don’t have to worry about instruction set level compatibility.
It is ported (I think by AMD). I use it daily. Biggest problem is that the people maintaining PyTorch aren't very interested as they don't use it themselves. Hence there are silly bugs in it. Users contribute fixes and they get closed down as "unsupported" or get plain misunderstood.
Wait, is that actually true? The 4090 is superior in raster to the 7900XTX I thought? Happy to be corrected!
ultimately dlss/frame gen is a hack that no one on a high end card would ever use.
ray tracing nvidia is king but they got the jump. that will normalize over the next generation or two.
The irony of those technologies is that they shine best on high end cards.
If you're on a low end card and say pushing the limit of a game and are upscaling 720p to 1080p, then dlss can't do miracles, there will be tons of defects due how little information about the scene the upscaler has. But if you're on a high end system trying to upscale 1440p to 4k then even not only is the upscaler less likely to make a mistake, it's also much less noticeable.
Same with framegen, on a low end system that's getting sub 60fps, the framegen not only adds noticeable input lag, but often makes noticeable defective frames in between, if your system was already pushing above 60 then not only is the input lag much more bearable, the possible defective frames stay on screen for so little that you'd have to be paying attention to it to notice them most of the time.
They're still useful for low end systems, they can prolong their lifespan quite a bit for current gen games, but as this tech becomes the norm, developers rely on them to hit performance targets and your card will be outdated just as fast as older cards used to.
You may not be running DLSS with a super high scaling factor but I almost guarantee DLSS or DLAA is going to be on for most games nowadays if you want to keep framerates stable on ultra.
You say this as if momentum isn't one of the most powerful forces in the universe. Many a better technology has gone extinct because a competitor simply had more momentum. It's difficult to catch up and it's difficult to stop.
Nvidia also is aware of this and not making mistakes that many others do by sitting back and relaxing. They are doing everything they can to not let AMD even begin to compete, but they are also very aware that hardware is far from the full package. 75% max, but probably less.
For me DLSS looks worse than native. It looks all blurry. Even DLAA looks blurrier than native to me.
Alan Wake 2 will use FSR or DLSS for AA even at native resolutions, with no other AA option. I expect more games will do this in the future.
This is completely false.
Why?
Because Nvidia's DX12/Vulkan-style path in the driver is absolute trash, while their DX11/OGL4-era path is very well tuned. Nvidia's "advantage" is entirely software, the emperor wears no clothes.
With games that Nvidia's money couldn't cripple games, the 7900XTX, the card that both costs less and uses less watts, beats the 4090 in 1080p, 1440p, and 4k targets.
Have you seen fully realtime path-traced Cyberpunk?
You know very little about emulation.
Gamers buy 4090 only because DLSS is significantly better than FSR.
Nvidia is pouring $$$ into protecting their advantage as anemic as it might be. They give free compute and gpu resources to universities etc to cement the dependence on cuda. When you’re a beginner or someone who needs to just “get shit done” you’ll go nvidia at the expense of the greater good.
AMD's ML / ROCm support for their consumer cards is still awful.
I expect support will eventually be extended to a generation or two prior GPUs, but nothing beyond that.
https://rocm.docs.amd.com/projects/radeon/en/latest/index.ht...
It's actually quite brilliant. It never occurred to me before that the resources we had access to at university weren't the most popular because that's what the library wanted. They were the most popular because the provider made them cheap/free to widen their moat over the competition.
It reminds me of Thomson Reuters launching Eikon, a superior alternative (imho) to Bloomberg terminals. Every business school has Bloomberg terminals. But few/none had Eikon, which is why nobody knows how to use one.
It's a pretty solid strategy, but naturally favors the current incumbent who has the $$$ to throw around.
There's huge momentum on the "user" side of the equation that just getting cheap hardware and passable performance means Nvidia losing dominance.
This description isn't fair. Nvidia was pouring money into CUDA for 16 years. AMD/Intel only woke up once ChatGPT launched and nvidia started printing money. It's not fair to call it "protecting their advantage" when nvidia invested into the ecosystem for years and AMD couldn't even bother to ship working examples with OpenCL. I'm not sure why choosing the company that decided to invest in the platform for years is at the expense of the "greater good." What "greater good" is there above actually working software?
If they would have split resources off of Zen to spend on CUDA they would have failed at both and AMD would be a bankrupt husk.
The thing AMD doesn't have is the software stack.
You're right, it's harder. Saying this as someone who's done more work on the former than the latter. (I have, with a team, built a rocket engine. And not your school or backyard project size, but nozzle bigger than your face kind. I've also written CUDA kernels and boy is there a big learning curve to the latter that you gotta fundamentally rethink how you view a problem. It's unquestionable why CUDA devs are paid so much. Really it's only questionable why they aren't paid more)
I know it is easy to think this problem is easy, it really looks that way. But there's an incredible amount of optimization that goes into all of this and that's what's really hard. You aren't going to get away with just N for loops for a tensor rank N. You got to chop the data up, be intelligent about it, manage memory, how you load memory, handle many data types, take into consideration different results for different FMA operations, and a whole lot more. There's a whole lot of non-obvious things that result in high optimization (maybe obvious __after__ the fact, but that's not truthfully "obvious"). The thing is, the space is so well researched and implemented that you can't get away with naive implementations, you have to be on the bleeding edge.
Then you have to do that and make it reasonably usable for the programmer too, abstracting away all of that. Cuda also has a huge head start and momentum is not a force to be reckoned with (pun intended).
Look at TensorRT[0]. The software isn't even complete and it still isn't going to cover all neural networks on all GPUs. I've had stuff work on a V100 and H100 but not an A100, then later get fixed. They even have the "Apple Advantage" in that they have control of the hardware. I'm not certain AMD will have the same advantage. We talk a lot about the difficulties of being first mover, but I think we can also recognize that momentum is an advantage of being first mover. And it isn't one to scoff at.
I have a specific reason to believe this is all already optimized: poor adoption of Infinity Fabric by unsophisticated people, whereas NVLink is on every 3090 and DGX/HGX box. But that doesn't mean DC users haven't optimized.
1) the user above me thinks ML is just matrix multiplies, which it is far from that.
2) It's more than hardware, and one team has (huge) momentum
3) You can have better tech and not win. It actually happens relatively frequently.
There may be nothing better for the ML space than AMD being a meaningful competitor. It will give all of us better tools. It'll make AMD better and it'll even make Nvidia better, because (healthy) competition is good for everyone. But let's not be naive in thinking it is just hardware and that one example is enough. It's a good sign, but a Tesla driving on the highway in nominal conditions isn't even close to fully self driving. It's still a year away as it's been for a decade.
So I do hope this ages poorly. But I'm not holding my breath because as AMD succeeding is one of the best things for ML, hype is the absolute worst and there's way too fucking much of it. And it it takes 5+ years to age, well I wouldn't say it aged poorly.
Unfortunately AMD only claim parity with Nvidia for training right now, and uplifts on inference. That's the wrong way round to get people to spend time porting as a priority.
Anybody who has worked with them knows that it's basically the same.
Optimizing kernels for all architectures and all models is a hard problem as there are lot of cases to handle, but getting good utilization for a particular model on a particular architecture (like MI300 / Mixtral) is not that hard.
OP and OP’s OP never frame it about that. But you can still get the answer you are looking for from their answers.
Coming from FPGA (both Verilog and VHDL), to me CUDA offers the best way to handle all three parts in comparison to let‘s say OpenCL, Vulkan, or Metal.
For game developers this might be different, as here you always use high-level language and packages (in most cases).
AMD need‘s to rethink there strategy regarding compute tool stack radically. If the HelloWorldMatrixMul is longer as ten lines and not compile on every new AMD powered notebook out of the box, they have no chance to beat NVIDIA. (I know, that the last point is not given for NVIDIA. But it would give AMD the doubleplus good, currently missing.)
... I'm not sure what you mean that you always use high-level language and packages... But it doesn't sound right.
The reason game developers use DX/VK/GL as opposed to cuda is entirely: A. You want access to non-nvidia customers B. You need access to more aspects of the hardware then is exposed to Cuda. Cuda provides a nice interface to Nvidia's compute. Almost no access to the entire rest of the fixed function pipeline.
Most operations are memory bound: slow because memory access takes longer than computation itself. The problem is we have to shove the whole model, billions of weights, in the SRAM once for every token generated. That creates slowness.
Something doesn’t quite add up for this dependence on CUDA. Sure, I’m installing some dependency hellscape Python research project from GitHub I reasonably expect idiosyncratic hardware dependencies, but surely the companies in this space are beyond that on the maturity timeline.
I think part of the problem is something you're touching on here. The ML software space is horrifically put together. It's a bunch of research projects cobbled together. Hence the dependency hell, and very few people who actually know how to port it. Cuda siits at the bottom of the stack with its design assumptions baked into everything above it (see PyTorch for a good example of that). Nobody wrote a hardware agnostic layer over the top, and now the thought of pulling out a Jenga block at the bottom scares everyone.
Cuda is an intentionally leaky abstraction layer over the hardware, and its worked.
That's before getting to any potential 8 vs 16 bit processing speed differences the two companies are arguing over here.
In any case, if the comparison was being done in 8 bit, the relative performance between the cards should be roughly similar as it is with 16 bit.
Its also important to distinguish 8 bit compute from the actual quantization scheme. vLLM in fact supports 4 bit AWQ, but the computation is not done in 4 bits on the GPU, and it wouldn't even be close to pure 4 bit inference even if the hardware and vLLM supported that.
Also, compute stack is too rough right now. Not sure ROCm will be comparable with CUDA in terms of usability in the next couple of years at least.