QUIK is a method for quantizing LLM post-training weights to 4 bit precision
github.com
github.com
[1]: https://www.intel.com/content/www/us/en/docs/intrinsics-guid...
Fair point. It might help if the system is DRAM bandwidth limited, so reducing the data size helps even though individual operations take multiple instructions. But that is not the situation with todays hardware.
Intuition says, yes. Would appreciate some practitioner/theorist on the subject to say what is the impact of lower precision on accuracy/output of a model.
Here's a reddit post showing the 2.5 (exllamav2) quant as incredibly bad, at least: https://www.reddit.com/r/LocalLLaMA/comments/16mif47/compari...
There are techniques like "quantization-aware training" [1] which aim to reduce the impact.
However, if your benchmark is "output quality for a fixed amount of compute" (e.g. running real time voice recognition on a smartphone) a larger model that's been quantized might perform better than a smaller model that hasn't.
[1] https://www.tensorflow.org/model_optimization/guide/quantiza...