625 karma · joined November 3, 2012
My Ph.D. projects:
www.github.com/Maratyszcza/NNPACK
www.github.com/Maratyszcza/PeachPy
www.github.com/Maratyszcza/Opcodes
Cool demos:
www.github.com/Maratyszcza/blis-bench
www.github.com/Maratyszcza/laff-demos
Note: I work for Google, but speak for myself.
CPU is the default backend in TensorFlow Lite, and CPU inference always works and produce correct result. GPU/DSP/NPU inference can be faster, particularly for large models on high-end SoCs, but generally you need to make sure that the model is supported on the IP block, the result is correct and performance is better than the CPU baseline. And that quickly gets very complicated:
1. NN API, and TFLite GPU/DSP backends support a limited subset of all TensorFlow Lite operators, and if a model is only partially offloaded to GPU/DSP/NPU, part of it will still run on CPU, and commonly synchronization overhead kills all potential speedups of the specialized hardware. The situation is even worse in CoreML, as CoreML doesn't provide an API to even learn which operators failed to offload to GPU/NPU.
2. Bugs in GPU shader compilers and NN API drivers do happen, and unless your model is a standard MobileNet, you're likely to hit them at least on some mobile phones. Then you'd need an infrastructure to detect this situation and disable offloading the model to this IP block on particular phones.
3. Low-end SoCs usually completely lack DSP and NPU, and their GPU is often slower than CPU even in nominal peak performance. This happens because CPU cores in low-end SoCs are typically just downclocked versions of the CPU cores in high-end SoCs, but low-end GPUs have 8-16 times fewer GPU cores than their high-end counterparts.
Compute shader extension for WebGL 2.0 would be cool, but it would require to port a large part of OpenGL ES 3.1: OpenGL ES 3.0 / WebGL 2.0 doesn't include even random access buffers (SSBOs)
The source code is available on GitHub [3].
[1] https://maratyszcza.github.io/laff-demos/dgemm.html
[2] https://www.edx.org/course/linear-algebra-foundations-fronti...
Implementation details of course differ between frameworks, but luckily neural networks are very robust to noise. In my experience, changes like using Winograd/Fourier for convolutions, or even running the whole thing in FP16 do not result in noticeable artifacts, and these are among the biggest differences you could have between frameworks.
Disclaimer: I work on Caffe2 team (not on ONNX, though)
import opcodes.x86_64
isa = opcodes.x86_64.read_instruction_set()
print(sum(len(instruction.forms) for instruction in isa))
>>> 6020
As an another example, here is number of instruction forms (e.g. mnemonic name + operand types) over time on Intel CPUs: http://imgur.com/a/AVPcq