Demystifying NPUs: Questions and Answers
thechipletter.substack.com
thechipletter.substack.com
Normally I think of something like CUDA to get code running on a GPU. Can code targeting a CPU or GPU automatically be sped up running on a NPU? Or does code explicitly need to target the NPU?
Are they not basically identical hardware?
May be a degree of software compatibility at the highest level - eg PyTorch - but the underlying software will be very different.
A NPU does strictly only the operations required for ML inference, which use data types with low precision, i.e. 16-bit or 8-bit types.
A GPU is optimised for 3D rendering (and is useful for parallel computations in general). An NPU is optimised for neural network inferencing. These algorithms both involve matrix mathematics but they are not the same. The NPU hardware design matches the deep neural network inferencing algorithm. For example it has an "Activation Function" block dedicated to computing the activation function between neural network layers. It is optimised and specialised for one very specific algorithm: inferencing. A GPU would beat an NPU for training, and any other parallel computations besides inferencing.
I wonder if NPU will supercede GPU as in General Processing Unit now that it has finally entered the wider lexicon, relegating GPU back to Graphics Processing Unit or video cards.
And no, GPGPU (General Purpose Graphics Processing Unit) is a bloody stupid term to be bluntly honest.
No, GPU is almost universally used to mean GPU. There is nothing graphical about cryptocurrency, "AI" (sans image generation), protein crunching, and whatever else they are being used for that aren't graphical.
I question how many people are even aware the G is supposed to stand for Graphics anymore. The nomenclature is outdated and doesn't reflect reality anymore.
Most people aren't aware, because the industry never called attention to the morphng definition of GPU.
At that, CPUs aren't Central (for large-scale array-oriented computing workloads) anymore either. (They still are for enterprise or web workloads. Or, they're "Central" in terms of coordinating GPUs and moving data around. But no longer "Central" in terms of "does most of the computing".)
Very rarely, an acronym is “retconned” into a more appropriate expansion or simply starts being considered a regular word not standing for anything.
I’d strongly challenge your assertion that this has happened for “GPU”.
If you are offended by the term GPGPU, maybe we could use the name Compute Processing Unit :-).
I can certainly drink to that.
https://developer.nvidia.com/blog/programming-tensor-cores-c...
For example Apple’s NPU can’t do FP32 precision, it can only do FP16 and less.
Back in reality, that's not how any vendor intends their NPUs to be used. They provide high-level libraries that implement ONNX, CoreML, DirectML, or whatever and they expect you to just use those.
Will be interesting to see if other high level accelerator supporting languages like Chapel or Futhark or JAX end up getting NPU backends, it might give them a nice boost over the proprietary C++ inspired language.
Edit: JAX has TPU support.
I haven't done an in depth look but most matrix math accelerators (eg AMD, Intel and Apple) seem to provide C/C++/Python APIs for describing the computations but the code executing on the NPU is not compiled from user C++ code.
Apparently eg in Intel's stuff there's a custom run-time compiler consuming this kind of IR (intermediate representation) in the accelerator sw stack: https://docs.openvino.ai/2022.3/openvino_ir.html & https://github.com/intel/linux-npu-driver
And on AMD from user POV it doesn't seem too different: https://ryzenai.docs.amd.com/en/latest/devflow.html
"Will be interesting to see if other high level accelerator supporting languages like Chapel or Futhark or JAX end up getting NPU backends, it might give them a nice boost over the proprietary C++ inspired language."
As you say, the GPU (or NPU, TPU,..) don't run C++ or anything derived from it. The "runtime" (~backend) will usually emit some kind of hardware dependent format (or again, an IR) like SPIR-V, PTX, etc.
But the backend itself is usually written in C++ (due to performance reasons), and there is really no way to get around that. Interacting with that from Python (or Jax) is a usability win, but there is zero difference in functionality. I.e. there is no proprietary C++ inspired language in play here. Hence no way to get a boost.
In the Jax style implementation scenario the compiler part of JAX is better inspiration, maybe along the lines of this case study of a path tracer running on a TPU: https://blog.evjang.com/2019/11/jaxpt.html - I don't think Chapel or Futhark would adopt the same approach as such but it's at least some kind of existence proof of a compiler targeting it from a high level language for a non-machine learning code.
The main reason for this lack of direct programmability is that NPUs are fast-evolving, optimized technology. Hiding the low-level interface allows the designer to change the hardware implementation without affecting end-user software. For example, some NPUs can only work with specific data formats or layer types. Early NPUs were very simple convolution engines based on DSPs; newer designs also have built-in support for common activation functions, normalization, and quantization.
Maybe one day, these things will mature enough to have a standard programming interface. I am skeptical about this becoming a reality any time soon. Some companies (like Tenstorrent) are specifically working on open architectures that will be directly programmable, I'm not sure whether their approach translates to the embedded NPUs, though. What would be nice is an open graph-based API and a model format for specifying and encoding ML models.