Accelerating Neural Networks with Binary Arithmetic
software.intel.com
software.intel.com
One thing that looks attractive to me is that FPGAs can be used very effectively for this purpose because these will take up much less LUTs and nets for specific purpose can be manually programmed in HDL.
They are also much better than GPUs for in-house deployment because they consume a lot less power which in turn reduces cooling power. They also don't get 'refreshed' every year which can be a good or bad thing because you don't have to upgrade so frequently and possibly lag being on performance compared to rest of the industry.
FPGAs have often been acknowledged as being a low power solution for the same precision but they're tremendously difficult to program. Also when you want FP operations on FPGA you're using up a LOT of chip space, because they were meant for GPUs.
So this kind of light weight neural nets will be much more suited for FPGAs. I've tried some basic nets but they are better for GPUs since you can fit larger nets without problem and RAM is another problem with FPGAs.
I'm thinking that these nets would be perfect since they take up much less space on the FPGA and consume a lot less RAM. (The lower RAM consumption is talked about here[1]) Also, the simple construction of the net with basic operations would help the programming difficulty.
For example, see Deephi's presentation at HotChips.
The OP is only suggesting binarisation for the feedforward use of nets, not the learning and training phase.
from the OP : " Real valued gradients are required for SGD to work. The weights are stored in real valued accumulators and are binarized in each iteration for forward propagation and gradient computations. "