comments on the X post point out that 8B models is pretty good.
What type of VRAM would you need to run 8B locally?
What type of VRAM would you need to run 8B locally?
( Q / 8 ) * B = GB RAM, where Q is the Quant level and B is the model size.
So a Q4 7B model is ( 4 / 8 ) * 7 = 3.5GB RAM (or VRAM).
A non-Q model is 16, so 2 * B.
I believe this is before context, which adds a bit.
https://ollama.com/library/llama3/tags
Llama3 8B fp16 is 16GB, q8 is 8.5GB and q4 is between 4.3 and 4.9GB.
Llama3 70B fp16 is 141GB, q8 75GB and q4 is between 40 and 43GB.
I use the Q5_K_M GGUF, which is >99% the same as the original.
I've seen tests that there is close to no divergence with these quantisations, but it rises steeply going lower:
https://www.reddit.com/r/LocalLLaMA/comments/1816h1x/how_muc...