just a few weeks ago they backported a kernel oops amdgpu null dereference into stable, it's still not fixed.
just a few weeks ago they backported a kernel oops amdgpu null dereference into stable, it's still not fixed.
Whenever I rebuild llama.cpp, I wind up using the Vulkan build anyway.
They have made a few attempts at investing in the hardware, but the software side is letting them down hard, and that part is almost entirely their own fault. They have underinvestment in their own stack, and into popular standard and Community libraries that would make it easy to use their gear.
Also, when you do Python work, you have to remember to install dependencies like pytorch from an alternate repo, otherwise you wind up with the CUDA versions that only use the CPU and not the GPU through ROCm.
I haven't finished comprehensive tests but I found:
- Vulkan is up to 2.25X faster for most coalesced, strided and interleave variants for memory-side scheduling/access shapes
- 3.3X faster on specific dot-path sweeps, including for scalar-dequant
- For matched LDS, Vulkan can be 8-14X+ faster (!!!) than matched HIP LDS
HIP doesn't always win against RADV/ACO, but on dispatch/runtime, it does appear to be quite a bit faster than HIP/LLVM on gfx1151 (Strix Halo). I'll be publishing sharing full data once I also run vs gfx1100...
Personally I’ll never buy AMD again.