No, 4bit quantization is the typical case.
At 4bit you can fit twice the parameters of 8bit in the same space for far better performance/perplexity/quality.
Running LLMs higher than 4bit is atypical and almost always sub-optimal (compared to running a model half the size in 8bit).
Even pretraining and finetuning in 4bit is likely to become the norm soon as fp4 becomes more well understood.