Yeah, there's a tradeoff. pre-training with QAT needs more data and a ton of hardware, so it's easier to just take an existing open weight model and quantize it. This works, but it'll underperform a model that had quantization as a training target.
Matters more for 1-bit and 2-bit models.