Given how close it is to 405B in performance it would be interesting to see which has the edge comparing an unquantized 3.3-70B against 405B quantized to be the same size.
That would be 1.38 bits per weight on average, which I can confidently guess would not perform well.
BitNet is functional at 1.58 bpw.
The model card says the 70B is 16 bit so I think you have twice that
It's kind of amazing how there seems to be a wall where sizing up the model starts to diminish in terms of intelligence gains. I guess that's why we can still compete with whales even though their brains are like twice as big as ours.
There is a line of thinking that the required tokens shoved through training also needs to go up, perhaps super-linearly to the model size. And of course, there's a line of thinking that there are diminishing returns in all things, including these models. Both could be true!
I have tried 405B at 1-bit quantization. It remains coherent, but didn't seem to be any better than 3.1-70B.