Anecdotally, it seems most models can be quantized to 8 bits without much loss of accuracy, and fixed point arithmetic requires much less hardware. Training is still done with floating point though.
Here is another paper demonstrating very good results with just 6 bit gradients: https://arxiv.org/abs/1606.06160
https://arxiv.org/pdf/1502.02551.pdf
https://arxiv.org/pdf/1610.00324.pdf
I wanted to have some basic idea about hardware so I did some "research" (googling) and ended up giving a short informal talk. My slides with some links are here: