> 4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.
Doesn’t that describe the majority of quants on higgingface?
Doesn’t that describe the majority of quants on higgingface?
Matters more for 1-bit and 2-bit models.