Didn't this paper demonstrate that you only need 1.58 bits to be equivalent to 16 bits in performance?
That doesn't mean training with all integer math, but certain tricks are used to specifically plan for the end weight size. I.e. fake quantization nodes are inserted to simulate int4.