How to Run DeepSeek R1 Distilled Reasoning Models on RyzenAI and Radeon GPUs
guru3d.com
guru3d.com
The 32B model achieves about 25 tokens/s, which is faster than I can read. However, the "thinking" time is mostly a lower quality overhead taking ~1-4 minutes before the Solution/Answer
You can view the model performance within ollama using the command: /set verbose
The good thing of 32B is being as good as 70B at many benchmarks according to Deepseek documentation
https://huggingface.co/deepseek-ai/DeepSeek-R1#distilled-mod...
But I cannot find it in LM Studio, what am I doing wrong that I only find distilled models?
I didn't go into detail about how to setup openweb-ui, but there is documentation for the on the project's site.
Environmetn="ROCR_VISIBLE_DEVICES=1"
The 't' and 'n' are transposed.
Might not be appropriate for this model, but it could be for small models.
1. sudo pacman -S ollama-rocm
2. ollama serve
3. ollama run deepseek-r1:32b
I tried running a model larger than ram size and it loads some layers into the gpu but offloads to the cpu also. It's faster than cpu alone for me, but not by a lot.