GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
When the qwen 4 series is released I am hopeful there'll be a strong model with the same architecutre. as qwen3.8-flash-next.
| 0 N/A N/A 321801 C llama-server 16540MiB
| 1 N/A N/A 321801 C llama-server 22404MiB
| 2 N/A N/A 321801 C llama-server 23480MiB
| 4 N/A N/A 321801 C llama-server 43982MiB
| 5 N/A N/A 321801 C llama-server 41656MiB
| 6 N/A N/A 321801 C llama-server 22366MiB
| 7 N/A N/A 321801 C llama-server 16208MiB
That's 3.86 bits per word I could run the 4 bpw one as well (there is still one spare gpu and another 27G on the ones listed above. DS4 uses about the same memory, is a little bit faster (though I suspect that by the time GLM 5.3 support is a bit more mature the speed difference will have evaporated).From the homepage of the inference engine:
"DwarfStar aims to be the best way to run a few excellent large language models on consumer hardware (that is, hardware that people can actually own). To reach this goal, we are building a small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)."
The model has been around longer than that. DS3 = Deepseek 3 etc. I think Salvatore named it pretty cleverly but he doesn't automatically get to own a two letter acronym.