We don't current compile in CLBlast or ROCm support but if there's a lot of demand for this, we'll definitely add it in the future. One concern is not wanting to bloat out the binary size too much (CUDA is already huge!) but given how big the LLM models are anyway, maybe it's not a huge concern.