Title misses that it's "weights and activations" as in they (from my skim) are doing all the math with mostly 4-bit values. Normally quantized weights are converted to 32-bits (edit: or more generally a type for which there are native machine instructions to multiply/add, as mentioned below) as they are applied during a forward pass which saves memory but incurs extra processing. They are keeping everything in (mostly) 4-bits to make it faster.