Apple should've invested more in bandwidth, but it's Apple and has lost its visionary. Imagine having 512GB on M3 Ultra and not being able to load even a 70B model on it at decent context window.
qwen 2.5 coder 1.5b @ q4_k_m: 1.21 GB memory
qwen 2.5 coder 1.5b @ q8: 1.83 GB memory
I always assumed this to be the case (also because of the smaller download sizes) but never really thought about it.
> entire point...smaller download could not justify...
Q4_K_M has layers and layers of consensus and polling and surveying and A/B testing and benchmarking to show there's ~0 quality degradation. Built over a couple years.
Llama 3.3 already shows a degradation from Q5 to Q4.
As compression improves over the years, the effects of even Q5 quantization will begin to appear
That doesn’t necessarily translate to the full memory reduction because of interim compute tensors and KV cache, but those can also be quantized.
As for CPUs, Intel can only go down to FP16, so you’ll be doing some “unpacking”. But hopefully that is “on the fly” and not when you load the model into memory?