> - Training isn’t done at 4-bits, to date this small size has only been for inference.
Wasn't there a paper from Microsoft two weeks ago or so where they trained on log₂(3) bits?
Wasn't there a paper from Microsoft two weeks ago or so where they trained on log₂(3) bits?
This makes network minimise loss not only with regard to expected outcome but also minimises loss resulting from quantisation. With big networks their "knowledge" is encoded in relationships between weights, not in their absolute values so lower precision work well as long as network is big enough.