Aiter: AI Tensor Engine for ROCm
rocm.blogs.amd.com
rocm.blogs.amd.com
El Capitan is #1 in TOP500. Frontier is #2, LUMI is #8.
ROCm development is probably mainly driven by the needs of these supercompuers' users currently.
So, we're seeing the tip of the iceberg.
Also ROCm packages continue to land on Debian, so there's more than meets the eye.
Note: Search "AMD Instinct" at https://top500.org/lists/top500/list/2024/11/. There are way more systems.
Simulations run on FP64, and you have to since you're already approximating stuff with numerical algorithms (analytic solution of many things are impossible anyway). Even if you can do things with FP8, transferring everything to GPU is not trivially possible.
A simulation contains tons of different algorithms, and not all of them can be modeled as a set of matrix operations effectively. Also, moving kernels in an out of GPU is not an instant affair, plus moving data to GPU is always more expensive.
You have GPUDirect and MultiDMA engines in modern GPUs, but they need hardcore coding and knowing what you're doing if you're not solving popular stuff with established libraries and so on.
Plus, if you don't prefer to be vendor locked, at least one of the vendors artificially limit the performance you can get from their cards.
On the other hand, all of the prominent linear algebra libraries squeeze out the CPUs you have relatively easily, and you don't have to have matrices and vectors to get this performance from CPUs anyway.
Lastly, I want to touch on that parallelization such problems are not always trivial even on CPUs. When you go multinode via MPI, things get fun. Getting GPUs into that mix is somewhat of a madness if you're not prepared.
I've been volunteering with Debian to help package ROCm for four years now, but today it officially became my full-time job. AMA.
Man, I'm old. :)
As messy as ROCm's packaging is, I can't imagine spending all day every day trying to fix it.
I'll be doing things like creating new packages in main, helping to get support for the HIP language embedded into existing dpkg tooling, helping to get GPU architecture awareness integrated into the Debian CI infrastructure, helping to enable ROCm support in other libraries and applications packaged for Debian, and ensuring that everything in Debian is successfully imported into the Ubuntu universe repositories.
Integrating HIP support into Debian so that it feels as natural as C or C++ and 'just works' across dozens of GPUs is a job for more than one person. That is why I'm glad there have been so many volunteers in the community stepping forward to help with various pieces.
One of the issues I've had with ROCm is not so great support for commercial GPUs. This is specifically with RX 7XXX series. Do you think there is any chance it will improve in future?
AMD just drops your card within a few years it seems like and drops your card from the current releases. Makes me favor Nvidia.
The good news is that I have at least one AMD GPU of each architecture from Vega to RDNA 3 / CDNA 2 on the Debian ROCm CI. Debian Trixie has packages built and tested for every modern discrete AMD GPU from Vega to RDNA 3 / CDNA 2. (I'd have liked to include RDNA 4 / CDNA 3, but the effort was quite resource constrained and the packages are a bit old. I'm hoping to improve upon that going forward, but Trixie is already in feature freeze so it will have to wait for the next release.)
I personally own much of the equipment for the Debian ROCm CI and I can promise I will continue testing new releases on old hardware for a very long time.
It's perhaps worth mentioning that Framework has directly supported Debian in providing access to hardware with AMD NPUs and iGPUs. I'm typing this message on one of two Framework 13 laptops that they donated to support Debian in this effort. I will be using it both for testing gfx1103 support on Debian and for testing the NPU packages when they become available. Framework also generously offered to provide one of those desktop systems you linked for the Debian ROCm CI [2]. It would also be used as a CI worker for the NPU runtime libraries once those are packaged.
[1]: https://www.phoronix.com/review/linux-614-features [2]: https://ci.rocm.debian.net/
OS is Debian Trixie (Testing). No secret sauce. Install & go. Everything is working perfectly.
Seems like a problem since AMD wants to go after AI capex?
For example, high bandwidth, low latency interconnects, supporting GPU direct network messaging and IO, are important.
High memory bandwidth is also quite important.
Debugging and performance profiling at scale also commonly uses similar tools.
OTOH, if you don't think nobody is focusing on FP64, look YoY performance gains on both CPUs and FPUs for high precision floating point performance. You'll be surprised.
I guess it's vibecoding "AI"...
I think most people don't want to have to think about vendor lock-in related bullshit. Most people just want their model to run on whatever hardware they happen to have available, don't want to have to worry about whether or not future hardware purchases will be compatible, and don't want to have to rewrite everything in a different framework.
Most people fundamentally don't care about ROCm or CUDA or OneAPI or whatever else beyond a means to an end.
What do you mean? Having ROCm fused MoE and MLA kernels as a counterpart to kernels for CUDA is very useful. AMD needs to provide this if they want to keep AMD accelerators competitive with new models.
Upstreaming that might actually help researchers doing new stuff vs. the narrow demographic of people speeding LLMs on MI300X's.
Do you know what the RT in TensorRT stands for? hint: AITER has nothing to do with TensorRT.
In theory they might not even need to be involved in optimising compute kernels, there is probably some PhD student who'll do the work because they want to be a kernel-optimising specialist. In practice a few strategic applications of paid talent is all they really need to do. Everyone wants to diversify off Nvidia so there is a lot of interest in supporting AMD if they are willing to push out firmware that multiplies matrices without crashing. Which has been a weird sticking point for AMD for a surprising amount of time.
Back in the day you had to optimize your card for Quake, do everything to make it run well. Now you have to do that for Pytorch.
That is exactly the attitude that got AMD out in the cold away from the AI revolution; they learned a lot of stupid lessons about optimising to specific games and present-day use cases instead of trying to implement general capabilities to a higher standard like Nvidia did in CUDA. They ended up a decade away from a multi-trillion dollar market
PyTorch might be special. I wouldn't be at all surprised if AMD does have a dedicated engineer working on PyTorch. But their problem to date hasn't been that their engagement with PyTorch, but rather that literally nobody could make PyTorch work on AMD cards which had buggy and terrible support for GPGPU work. If they fixed that some random might do the work without their involvement because a lot of people want to see that happen.
Considering its importance, it shouldn't be one engineer. It should be 50+.
I think it’s ok that stuff is tried first in Torch extensions. That’s how Flash Attention started after all and the same is true for newer kernels in CUDA-land (fused MoE, MLA, Marlin, etc.).
With regards to TorchScript, that’s really legacy - torch.compile is where it’s at. This post seems to suggest that the kernels work with torch.compile: https://rocm.blogs.amd.com/artificial-intelligence/DeepSeekR...
They don't seem to care, or don't understand how to get broader adoption.
For some reason AMD's management is dead set on targeting only the high end part of the market. Like, for example, look at this blog post. Which model they're testing? DeepSeek R1, the 671B behemoth that no normal person can run. Or look at any of their tutorials/docs and see which GPUs they support - it's always only either the unobtanium-grade enterprise GPUs, or high end workstation cards that no one buys. And if your strategy is to target only the super rich entities then a little jank in the software isn't really all that punishing - if you can afford to drop a few million on GPUs then you can also afford to hire someone to spend a few weeks getting AMD's software to work/get it tuned by tweaking two dozen environment variables they do seem to like so much/etc.
Because those people are dropping $100 billion on GPU clusters and individuals are not
NVIDIA GPUs sell so well because they work with what researchers actually use.
from aiter.tuned_gemm import tgemm
import torch
class LinearLayer(torch.nn.Module):
def **init**(self, in_features, out_features):
super(LinearLayer, self).**init**()
self.weight = torch.nn.Parameter(torch.randn(out_features, in_features).cuda())
self.bias = torch.nn.Parameter(torch.randn(out_features).cuda())
def forward(self, input):
input = input.cuda()
return [tgemm.mm](http://tgemm.mm/)(input, self.weight, self.bias, None, None)it's __init__ of course
Also the tgemm.mm has to be a torch module (at first I thought this was some lowlevel library which they now have a preview of, because there is a ROCm-torch already ...) which is evident from the table just before the summary. That table also smells like they are mostly focused on inference...
EDIT: seems official ROCm-torch is also based on HIP.
I'm assuming the scrambled annotations were due to some odd chain of things the code went through on the way to becoming a post.
Maybe they did it as a parable about the problems of having many layers of abstraction causing processes with unintended consequences?
EDIT: They fixed the code pretty quickly
Also I assume nvidia does the same thing but it is still hilarious that this is how it works
https://github.com/ROCm/aiter/blob/main/aiter/configs/bf16_t...
See Shader ISA: https://www.techpowerup.com/gpu-specs/radeon-rx-7600-xt.c419...
In any case, the possible values can be found in the LLVM documentation [1]. I would recommend looking closely at the notes for the generic ISAs, as they highlight the differences between the ISAs (which is important when you're loading code built for one ISA onto a GPU that implements a different ISA).
Using one high level language and assembly sounds fine, but four feels incoherent. Would love to know why this has had happened.
"This infrastructure is built upon a variety of underlying technologies, including Triton, CK (Compute Kernel), ASM (Assembly), and HIP (Heterogeneous Interface for Portability)."
There is absolutely nothing out of the ordinary here. Yes, it's multiple languages, but not any more or any different than what you'd use on an Nvidia platform (except obviously for the assembly part -- AMD's ISA is different from PTX, but that's to be expected).
It's having both Triton and HIP in the same project which I find weird. It feels very fragmented to me to use two high level languages. Maybe it makes sense given Triton is easier to use but less fully featured, but it definitely didn't strike me as normal.
I would be interested to know if NVIDIA use more than CUDA and PTX/SASS to write CUDNN and CUBLAS.
But you are right CK is indeed a library, thanks for pointing that out.
On one hand, cool. On the other hand wow have they been leaving a lot of performance on the table.
How does the performance compare to NVidia now?
I have also run Hunyuan3d-2 and generated 3d models. You would've to separate out the model generation and texture generation phase, but it works.
I run ComfyUI and bootleg gguf models. This is all on windows. Now even WSL2 works, so I am using Ubuntu-24.04 on Windows 11 to run Hunyuan3D-2.
For LLMs, llama.cpp native binaries are available. Everything just works out of the box.
Aiter is a ROCm library.
ROCm is the thing that is like CUDA, but for AMD.