What about the 4070 which 'only' has 12GB of VRAM?
I'm not certain 30B models will fit completely on 12GB, even with this quantization.
Obviously that would be useless for games, but for LLMs it may be an option?
Llama.cpp allows you to specify that n layers load into the GPU's VRAM, and the rest load into main memory.