You don't need quantization aware training on larger models. 4 bit 70b and 405b models exhibit close to zero degradation in output with post training quantization[1][2].
[1]: https://arxiv.org/pdf/2409.11055v1 [2]: https://lmarena.ai/
[1]: https://arxiv.org/pdf/2409.11055v1 [2]: https://lmarena.ai/
Same reason why you can get a pretty good reconstruction when you add random noise to an image and then apply a binary threshold function to it. The more pixels there are, the more recognizable will be the B&W reconstruction.