./llama-bench -m /data/ai/models/llm/gguf/mistral-7b-instruct-v0.1.Q4_K_M.gguf
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 1 ROCm devices:
Device 0: AMD Radeon 780M, compute capability 11.0, VMM: no
| model | size | params | backend | ngl | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------------: | ---------------: |
| llama 7B Q4_K - Medium | 4.07 GiB | 7.24 B | ROCm | 99 | pp512 | 242.69 ± 0.99 |
| llama 7B Q4_K - Medium | 4.07 GiB | 7.24 B | ROCm | 99 | tg128 | 15.33 ± 0.03 |
build: e11bd856 (3620)Let’s be honest, it might not be awful but it’s a nonstarter for encouraging local LLM adoption and most will prefer to pay to pay pennies for api access instead (friction aside).
However, if for anyone that is looking to use a local model on a chip with the Radeon 890M:
- look into implementing (or waiting for) NPU support - XDNA2's 50 TOPS should provide more raw compute than the 890M for tensor math (w/ Block FP16)
- use a smaller, more appropriate model for your use case (3B's or smaller can fulfill most simple requests) and of course will be faster
- don't use long conversations - when your conversations start they will have 0 context and no prefill; no waiting for context
- use `cache_prompt` for bs=1 interactive use you can save input/generations to cache
I don't think so. Humans scan for keywords very often. No body really reads every word. Faster than reading speed inference is definitely beneficial.
But to your other point, very little of the current popular ML stack does more than CUDA and MPS. Some will do rocm but I don’t know if the AMD iGPUs are guaranteed to support it? There’s not much for Intel GPUs.
I'm hoping that in combination with the gfx11-generic ISA introduced in LLVM 18, this will make it straightforward to enable compute applications on both Phoenix and Strix (even if they are not officially supported by ROCm).
[1]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
However the commits do have some caveats called out, as do the techniques they use to achieve the higher allocations.
As an aside, it just feels like a lot of hacks that AMD and Intel should have handled ages ago for their iGPUs instead of letting them languish.
If you only care about inference, llama.cpp supports Vulkan on any iGPU with Vulkan drivers. On my laptop with crap bios that does not allow changing any video ram settings, reserved "vram" is 2GB, but llama.cpp-vulkan can access 16GB of "vram" (half of physical ram). 16GB vram is sufficient to run any model that has even remotely practical execution speed on my bottom-of-the-line ryzen 3 3250U (Picasso/Raven 2); you can always offload some layers to CPU to run even larger.
(on Debian stable) Vulkan support:
apt install libvulkan1 mesa-vulkan-drivers vulkan-tools
Build deps for llama.cpp: apt install libshaderc-dev glslang-dev libvulkan-dev
Build llama.cpp with vulkan back-end: make clean (I added this, in case you previously built with a diff back-end)
make LLAMA_VULKAN=1
If more than one GPU:
When running, you have to set GGML_VK_VISIBLE_DEVICES to the indices of the devices you want e.g., export GGML_VK_VISIBLE_DEVICES=0,1,2
The indices correspond to the device order in vulkaninfo --summary.
By default llama.cpp will only use the first device it finds.llama.cpp-vulkan has worked really well, for me. But, per benchmarks from back when Vulkan support was first released, using the CUDA back-end was faster than the Vulkan back-end on NVIDIA GPUs. Probably same Rocm vs Vulkan on AMD too. But, zero non-free / binary blobs required for Vulkan, and Vulkan supports more devices (e.g., my iGPU is not supported by Rocm)-- haven't tried, but you can probably mix GPUs from diff manufacturers using Vulkan.