My understanding is that as a general rule:
( Q / 8 ) * B = GB RAM, where Q is the Quant level and B is the model size.
So a Q4 7B model is ( 4 / 8 ) * 7 = 3.5GB RAM (or VRAM).
A non-Q model is 16, so 2 * B.
I believe this is before context, which adds a bit.