oh yeah, in my experience anything below ROCm6.x really sucks.
I tried to run qwen2.5-32B on ROCm5.x and it was running at <15tok/s lol.
Have you tried running any sort of LLM inference on your MI25, or what NN workloads are you running?
I tried to run qwen2.5-32B on ROCm5.x and it was running at <15tok/s lol.
Have you tried running any sort of LLM inference on your MI25, or what NN workloads are you running?
No comments yet.