New Optimizations Improve Deep Learning Frameworks for CPUs
nextplatform.com
nextplatform.com
They have a 72x improvement, great! So, first start with a really crappy baseline not tuned for your high end 68 core machine - which runs slower than the 22 core machine so that you have a tonne of space to improve.
Then pump up the batch size to 2048 images per batch because it's an easy know to twiddle that has nothing to do with the implementation speed and you don't really care about the actual model accuracy and you just want to prove how fast you can make this. (I'm ignoring the recent work on making large batch training work, but the larger point is that batch size is a model hyperparameter that influences more than just speed)
And then ignore the fact that it's performance per $ that people care about, not pure speed.
Aside from the nonsensical performance numbers:
Theano, Neon, Torch, really?
No mention of PyTorch, Theano actually being discontinued..