QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
arxiv.org
arxiv.org
I need to seriously revise my definition of affordable commodity hardware
If you can't afford a hamburger, your problems are likely not in compressing trillion-parameter models.
I'm not in the field. Can someone explain how the sub-1-bit part works--are they also reducing the number of parameters as part of the compression?
It’ll be interesting to see if it works on the new mistral moe model, which is less sparse and probably trained more per param than these.
think of it like bog standard compression algorithm.