I hope Google will eventually reveal if the IPU shares any DNA with the TPU at all.
Now we are going to back to fixed function units as the calculus has changed in favour of power savings and against flexibility.
I'm interesting in seeing what advance in tech will change it again in favour of flexibility.
This has always been a trade-off. For instance even with desktops, once video handling became common we first had video decode/encode being handled by GPU acceleration. But now most CPUs include dedicated h264 (and more recently hevc) decode/encode support, eg [0].
Although it's much lower computation, hardware acceleration is also offered for audio, and I'd take a guess the reason the Apple APIs for checking which hardware supports hardware-based audio encode/decode [1] has been deprecated is now because all supported devices provide all possible capabilities.
[0]: https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video [1]: https://developer.apple.com/documentation/audiotoolbox/16204...
encoders: also much higher quality
https://www.nextplatform.com/2017/05/17/first-depth-look-goo... First In-Depth Look at Google's New Second-Generation TPU
I wonder if it's not so much about being better at image processing, but having control and direct access to the hardware. For the Adreno, you have to go through Qualcomm's drivers and are subject to their limitations (there is freedreno which works great, but I don't think Qualcomm allows it on handsets). You're also stuck with the sizes of GPU that are available on Snapdragon SoCs - if you want a larger one, too bad.
(Which isn't surprising; GPUs are really designed for 3D rendering and they're definitely overkill for simpler tasks. The general assumption is that you're fetching small locally-contiguous groups of pixels from main RAM, doing texture mapping and computations on them, then conditionally blitting the result to other locally-contiguous areas of RAM. Most of the infrastructure used for this is going to waste if you're just using it to do 2D image processing.)
I did find the original source to the Ars article [2] and it says that each IPU core has 512 ALUs. This seems more like a extremely wide GPU than the VPU. It also seems to be programmable in Halide [3], which presents a more "SIMT"-like interface, just like a GPU shader. It probably lacks the super fast bilinear texture mapping units that a GPU has, but otherwise seems very similar. It'd be interesting to know if the texture/pixel cache is handled automatically, or manually with a large register file like the VPU.
The article also features this quote, which also seems to back Google being unsatisfied with only high level access to the GPU:
>A key ingredient to the IPU’s efficiency is the tight coupling of hardware and software—our software controls many more details of the hardware than in a typical processor.
It would be really awesome if they opened this chip up to third party developers. Unfortunately the press release only mentions availability in the Camera API so far, relegating it to the same opaque blob status as all the other dedicated image processors :(
[1] https://github.com/hermanhermitage/videocoreiv/wiki/VideoCor...
[2] https://blog.google/products/pixel/pixel-visual-core-image-p...
It was mandatory for REDCINE-X Pro but since RED has implemented GPU support it’s also not really necessary.