Qualcomm’s Hexagon DSP, and Now, NPU
chipsandcheese.com
chipsandcheese.com
The HLOS (High-level OS) running on the Hexagon requires every "applet" to be signed by either the Qualcomm root cert or the OEMs cert. Usually, every phone has a set of generic Hexagon applets (or "skeleton libs") that are provided and signed by the OEM, which seem to be freely usable to offload some computational work to the DSP (mainly FastCV et al - https://developer.qualcomm.com/sites/default/files/docs/qual...). Those of course come with their own bugs: https://research.checkpoint.com/2021/pwn2own-qualcomm-dsp/
On some older SoCs, you were able to use a TOCTOU (Time of check to time of use) exploit to bypass the signature check by patching the applet loader shim in-memory, once it itself got authenticated: https://github.com/geohot/freethedsp/ (I have personally ported this to the msm8953, and it seems to work)
When I switched to NVidia I was surprised to find a much more open ecosystem with good public documentation. NVidia did have some tasty secret sauce stuff that they didn't expose outright, but they did what they could to empower developers to make the best use of the underlying hardware. They strike the right balance between openness and maintaining a competitive advantage, in my view.
Just my opinion based on working in both companies for a number of years. Thankfully I no longer have a dog in that fight.
> The HLOS (High-level OS) running on the Hexagon requires every "applet" to be signed by either the Qualcomm root cert or the OEMs cert
That's no longer true since quite some years now :) See the Unsigned PDs, which are allowed for general purpose compute since at least sm8150 (Snapdragon 855).
Note that the articles you mention says this about it:
> Signature-free dynamic shared objects are run inside an Unsigned PD, which is the user PD limited in its access to underlying DSP drivers and thread priorities. An Unsigned PD is designed to support only general computing applications.
I am becoming convinced that CPU (and maybe GPU) is the only viable accelerator on Android devices. All these fancy accelerators are just for phone makers to do their own thing (mainly camera crap). Might as well make it part of the ISP.
Also, I fear Apple is going to eat Android's lunch at this rate :(
I disagree the Hexagon is "hard" to program - the HVX intrinsics are fairly straightforward if you've used neon/avx/etc. The problem is that you don't get access to the HMX/HTA/(whatever they call the NPU currently), just HVX. So for running on DSP your options are (1.) Qualcomm's Hexagon SDK which doesn't give you access to the tensor hardware, and (2.) their SNPE SDK, which actually has access to HMX/HTA/etc, but doesn't let you program against the hardware directly. Instead, SNPE is supposed to convert your caffe/tf/onyx/torch models to Qualcomm proprietary format so it can map ops to the most appropriate hardware. What actually happens is that their conversion tools fall apart the first time you don't use one of their AlexNet/ResNet vanilla examples and try to convert a prod-grade model. Combine the brittleness of their conversion tools with the lack of documentation/support and it becomes impossible to use the hardware.
All that being said, would be interested to hear if anyone has had luck with other means of using their (admittedly impressive) hardware. Maybe a better approach would be to try using Halide for Hexagon and OpenCl for Adreno? They were also going to release a QNN update this year or next year - not sure if they've allowed anyone access to it. Its really a shame that the hardware is so good, but the software is so clunky.
The difficulty comes in launching your code on the device from the apps core. It's a bit unlike many other typical conventions for this.
I agree that once you're executing code there that "HVX intrinsics are fairly straightforward" etc
It's true, this is a big drawback. Something like OpenCL would be really nice. I haven't used OpenCL for a while but way back when the design seemed to consider that the memory used by the accelerator device was a separate destination to be copied to. For the SoCs where Hexagon shows up, that's generally not the case. But hopefully there's already some tweaks to OpenCL to enable different SMMU contexts instead of copies.
The article also complains about VLIW in the same paragraph, but I don't think VLIW makes things harder, it just makes problems more obvious. If you write ARM or x86 code that has dependencies between every instruction, that's going to suck too, you just won't know it until you run it, but VLIW will make it obvious if you just look at the generated code. For the kinds of programs that make sense to run on a processor like Hexagon, VLIW is fine.
The whole Hexagon environment is just so much better than any of the other similar DSPs I'm aware of: you can use open source LLVM to compile code for it (so you aren't stuck with an old version of GCC), and the OS is much closer to standard (e.g. thread synchronization is just pthreads).
I did a bunch of work on Hexagon and I like it a lot. It is my favorite in its class.
And there's Halide support too.
Hexagon is a difficult architecture to write code for but the benefits are worth it: it's the secret sauce for why Qualcomm's modems are so good. I see people getting all excited over AVX512 and I just think "well we had 2048 bit vectors years ago"
That’s ultimately the downfall of NPUs. If it’s not accessible it may as well not exist.
While Hexagon and HVX are publicly documented ISAs, HTA (legacy), HMX and friends are not publicly documented ISA extensions.