>I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality
There's literature on this actually, if you want to get that low the model really should have quantization as a pretraining target otherwise larger models that are quantized after collapse.
In fact, to do this effectively, you need somewhere near 50x chinchilla to do QAT on super tiny targets like 1-bit or even 2-bit. These large models already need an astronomical amount of training data that doing proper QAT that small isn't really feasible unless there are breakthroughs in the architecture.
4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.