>LLaMA 13b (24.5GB) and Ministral 7b (13.6GB)
But the HW requirements state 8GB of VRAM. How do those models fit in that?
But the HW requirements state 8GB of VRAM. How do those models fit in that?
If it means that weights for an LLM can be 4 bits well that's just mind boggling.
I was skeptical of it for some time, but it seems to work because individual parameters don’t encode much information. The knowledge is embedded thanks to having a massive number of low bit parameters.