Now all we need is better support for AMD gpus, both CDNA and RDNA types
(PyTorch does also support ROCm generally, it shows up as a CUDA device.)
However from experience with an AMD Strix Halo, a couple of caveats: it's drastically slower than Ollama (tested over a few weeks, always using the official AMD vLLM nightly releases), and not all GPUs were supported for all models (but that has been fixed).
If you want more performance, you could try running llama.cpp directly or use the prebuilt lemonade nightlies.