It's a nontrivial problem to do faster arithmetic operations with quantized values (which are not just rounded to fit in 4 bits but are usually block quantized with one or more scaling parameters stored at full precision) faster than to convert them to floats that the cpu/GPU is already optimized to multiply/add. That is what this project is proposing a solution to.