Making AMD GPUs competitive for LLM inference
blog.mlc.ai
blog.mlc.ai
There are two points I personally wanted to make through this project:
1) With a sufficiently optimized software stack, AMD GPUs can be sufficiently cost-efficient to use in LLM serving; 2) ML compilation (MLC) techniques, through its underlying TVM Unity software stack, are the best fit in terms of cross-hardware generalizable performance optimizations, quickly delivering time-to-market values, etc.
So far, to the best of our knowledge, MLC LLM delivers the best performance across NVIDIA and AMD GPUs in single-batch inference on quantized models, and batched/distributed inference is on the horizon too.
The catch is:
- MLC's quantization is somewhat different (though I havent run any perplexity tests yet)
- There is no CPU offloading (or splitting onto an IGP) like Llama.cpp yet (unless its new and I missed it).
Regarding quantization, we wanted to develop a code path that absorbs any quantization formats, for example, those from GGML or GPTQ, so that they could be all used. ML compilation (MLC) is agnostic to any quantization formats, but we just haven't exposed such abstractions yet.
On CPU offloading, imagine if you are writing PyTorch, it should be as simple as a one-liner `some_tensor.cpu()` to bring something down to host memory, and `some_tensor.cuda()` to get it back to CUDA - seems a low-hanging fruit but it's not implemented yet in MLC LLM :( Lots of stuff to do and we should make this happen soon.
I don't know whether there's a LLM inference benchmark in the CI suite, if not perhaps something like this should be included in it.
If you're building the stack from source or found it in a Linux repo, decent odds it'll work for you. More likely to work on gfx9 or gfx10 than the older cards. I think that's roughly the last five years.
If you use the official distribution, some parts are compiled to gpu-specific machine code and if your gpu isn't one of those, you can't use that library. I think there's a reluctance to compile the libs for GPUs that aren't in the internal CI in case they don't work.
As an anecdote, I do most development on unsupported hardware, unsupported distro and unsupported kernel, with the upstream driver, using whatever was on llvm main that morning. That mostly works despite positioning myself as most likely to run into bugs.
It's a mature technology used my millions of people every day.
Unlike GPGPU compute, for videogames performance directly affects usability.
For these reasons, the software on all levels of the stack might be more optimized.
Can you comment on how difficult it was to achieve this, and what the relative advantages b/w cards ? AFAIR, AMD cards were not not deemed competitive with Nvidia in DL space largely because of the amazing job Nvidia pulled off with CUDNN and its conv. kernels.
LLMs etc. OTOH doesn't really depend on convolutions (atleast the pure transformer bits), and instead depends a lot more on plain old GEMM + low-bit float/int compute.
Thanks for asking! I personally believe TVM Unity is a proper software stack for ML compilation (MLC), and its existing optimizations (e.g. TensorCore offloading) can be transparently transferred to AMD/Intel/Apple/mobile GPUs without too much engineering effort.
Of course my claim is limited to ML workloads. Not an expert outside the ML world, so I couldn't say for general HPC.
Good work though. And you have an activity community on github, congratulations.
Btw - I got biased sampling working in ad-llama! Catching up to guidance slowly but surely :)
Support in TVM’s graph IR (Relax) - https://github.com/apache/tvm/pull/15447 Support in TVM’s loop IR (TensorIR) - https://github.com/apache/tvm/pull/14862 Distributed dialect of TVM’s graph IR for multi-node (GSPMD-type): https://github.com/apache/tvm/pull/15289
The first target will be LLM's on multiple NVIDIA GPUs but as with all of MLC-LLM effort, the approach will generalize to other hardware including AMD's wonderful hardware.
One question: given your experience, when would you predict a near parity in software stack support between te different platforms, so that a choice of GPU becomes one mostly of price/performance? It does not need to be like the AMD/Intel in the CPU market where a consumer will have no doubts about software compatibility, but let's say like the gaming gpu market where a game having problems on a gpu architecture is a newsworthy exception that is quickly corrected.
might be interesting to team up
Oh great. The AMD RX 580 was released in April 2018. AMD had already dropped ROCm support for it by 2021. They only supported the card for 3 years. 3 years. It's so lame it's bordering on fraudulent, even if not legally fraud. Keep this in mind when reading this news. The support won't last long, especially if you don't buy at launch. Then you'll be stuck in the dependency hell that is trying to use old drivers/stack.
Both rocm and vulkan are supported in MLC LLM as mentioned in our blog post. we are aware that rocm is not sufficient to cover consumer hardwares, and in this case vulkan is a nice backup!
1: https://rocm.docs.amd.com/en/latest/release/windows_support....
But you seem dead set that there are no uses for ROCm so I'll leave you there.
That said, people are finding hacks...
https://old.reddit.com/r/Amd/comments/15t0lsm/i_turned_a_95_...
CUDA also works on consumer NVidia cards, not just business ones.
But my attempts to get direct ROCm support were thwarted by AMD.
By the way, still waiting for you to take me up on your 'bet'.
Now it's your turn Mr. "You're not going to find rx580's with enough vram for AI. Typically 4-8gb." This is completely false. Rather than acknowledging that you then tried to move the goalposts (much like I did in that past thread saying, "Oh, but maybe it's just my region where they don't.") It looks like we both behave a bit silly when trying to save face when we're wrong.
It isn't completely false. You're doing super limited stuff as a hobbyist that barely works.
> from your ignorant perspective
no need for the ad hominem.
is the output better such that it's desirable, or is this just a case of "too much performance hit for a marginal gain"?
Do you even know how brand recognition works?
The amount of people swearing off of AMD because of bad drivers ten years ago easily cost them a billion dollars. More than the cost of developing a good driver.
Not is not what I'm saying, I'm saying that if I buy up a bunch of rx580 cards, nobody is going to rent them from me.
Now, if I offered a bunch of MI250's on an hourly rate, you can absolutely bet people will rent them all.
usually at least some of the compute resources are "prosumer" workstations using commercial cards.
It was actually Apr 18, 2017 -- https://en.wikipedia.org/wiki/Radeon_500_series
The manufacturer can smugly proclaim they offered six years of support for a product that was on the shelf four years into the driver's lifecycle.
I'm in the HPC space, and pretty much everything I do on the GPU is bound by how quickly I can get data in and out of DRAM.
The point at which data motion to/from DRAM is not the bottleneck is when you do enough work per byte of data moved. How much work is that? On today's server GPUs it's in the region of 50--100 double precision floating point operations per byte. You can work out an exact number by taking the theoretical maximum floating point operations per unit time you can execute and divide by DRAM throughput (data moved per unit time).
O(50--100) double precision flops per byte is a _lot_ of work. We're talking BLAS-3 type operations. Anything level 2 or lower, or sparse operations, are typically bandwidth bound.
The more RAM you have on device, the fewer swaps you need to do (none at all if it's big enough), and the less those operations get amortized, bringing you closer to theoretical max throughput.
Matrix multiples are such that going from fitting 75% of your values to fitting 100% of your values can mean an order of magnitude speedup.
I work with problems so huge they do not fit on a single device. Multiple devices each own a small piece of the global problem. They solve a local problem and they must communicate over a network in order to solve the global problem. This is almost universally true.
You would know more than I in this field, and I expect it really is better to swap a partition than it is to use a network.
There are certain methods in HPC applications that are almost universally avoided because of how terribly they scale to a large distributed memory system. Matrix multiplies are one of them. Outside of a handful of ab initio computational chemistry algorithms (which are incredibly important), basically the only reason someone does a large dense matrix-multiply on a supercomputer is usually because they're running a benchmark and they're not solving a real science problem.
Folks more knowledgeable than me here feel free to jump in.
What I'm saying is it's not like a global solver (ie taking into account all to all interactions) wouldn't be more accurate right? It's just an insane proposition because surprise surprise that would require an enormous matmul during the update, which you can't do efficiently, even on a GPU, for the same reason the ML folks can't: the arithmetic intensity isn't high enough and so you can incur i/o costs (memory or network, same thing at this scale).
Neural networks, which are the basis for nearly all modern AI, are implemented as a mixture of sparse and dense matrix multiplies, depending on the neural architecture.
RAM bandwidth is basically always going to be a bottleneck. Microsft's proposal to get around this is to just keep everything in SRAM and pipe chips together.
As he went through the hard work of debugging the driver and getting AMD to care about it, I was expecting to see him in the acknowledgements.
Gfx drivers are complicated, long term projects. They're not something that get fixed with a phone call.
He found a bug, and made a bunch of noise about it. No need to make it more than it was.
I've worked on public software before, and it's left me with an extremely low opinion of this kind of behaviour, and the kind of people that show it.
BTW - that's exactly the dillema I'm pondering now when thinking about a new PC build to play with some machine learning / LLMs. In my personal situation, 7900XTX and 3090Ti are about the same price / TDP / profile. Almost the only difference is then CUDA vs ROCm and the fact that one is new and might therefore give me some basic warranty.
They’re not fast but they do share memory with the CPU.
The most interesting at the moment being the AMD Ryzen™ 9 7940HS Processor, 8 Cores/16 Threads.