For a 4x boost, I imagine Tensorflow and Torch will get support shortly after the GPUs start shipping in real quantities.
INT8 is a 4x for inference, but most people aren't using GPUs for inference atm.
https://blogs.aws.amazon.com/bigdata/post/TxGEL8IJ0CAXTK/Gen...
Unrelated question though: any chance you will do blog post/paper about how DSSTNE does automatic model parallelism and gets good sparse performance compared to cuSparse/etc?