Full model or a 4-bit quant? I have a 5090 and I'm not sure whether I should use a quant that fits within the VRAM or a much bigger version where I'd have to offload a lot to 64GB RAM and a beefy CPU (but still a CPU)
[0]: https://www.reddit.com/r/LocalLLaMA/comments/1vef79c/quantiz...