How to run 1.58bit DeepSeek R1 with Open WebUI
docs.openwebui.com
docs.openwebui.com
There is a PR where we offload non MoE layers to disk if you don't have enough RAM / VRAM, and it should make stuff faster!
1. GPUs are most likely at their limit in terms of FLOPs - float4 / FP4 is most likely the "final" low precision data-type. NVIDIA might provide 1.58bit support or FP2, but unlikely. If there was FP2, it might make it 1.5x faster.
2. Shrinking transistors might still have some room to go, but don't expect 2x or 4x faster - the majority of speedups was in tensor cores and low bit representation.
3. We might get more interesting transformer archs since DeepSeek showcased their unique arch for R1 / V3.
Due to these, I would first wait and see - ie I would actually wait until the next OSS model release say Llama 3, Gemma 3, etc and see what the large model labs are focusing on, then maybe I would wait for RTX 50 Super or a cheaper version or even RTX 60x series. Larger VRAM is always better.
SSD is good, but not that important - RAM is more important.
The goal was to showcase that MoEs quantized down to 1.58bit without any further training does in fact work!