As you said: everything works on llama.cpp
Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performance of models on Strix Halo, that is open for half a year (https://github.com/vllm-project/vllm/issues/34579#issuecomme...) and nothing is happening there. They do not care about those use cases. Seems like they are going with bit players that will run vllm inside datacenters racks. Hobbyists does not matter.