Because a SoC with limited CPU-core resources can't do everything in software, the chip contains many system components (hence System-on-Chip or System-on-a-Chip) that handle things the CPU cores then no longer need to do.
Think protocol handling or memory; instead of spending many clock cycles on handling the USB bus, you can leave that to the USB controller and only deal with what is actually relevant to your USB device. Same with the VOP (Video Output Processor) block, instead of spending many clock cycles on putting the right bits in the frame buffer you tell the VOP that you'd like the background to be orange and then only spend time setting the right bits for black text (for example). So instead of having to deal with many millions of bits, you only have to deal with less than 1% of them because every bit you don't set becomes orange. For other things like I2C, I2S, DMA, networking, cryptography, SD-IO, GPIO, PWM etc. the same applies. Instead of constantly spending time setting the right bit at the right time many times each second, you just tell a dedicated block on the SoC to do a thing in a certain pattern and it will do it for you, consuming on CPU core resources.
This also allows slow CPU cores that wouldn't be able to decode video in real time to offload the entire decoding to a video decoder block, and then tell the GPU part of the SoC that you're drawing a green rectangle somewhere and that's where it has to put the decoded video frames. Why would one do all of this? Because it's cheap and power-efficient, and that's how you make a big pile of money.
I have no idea how accurate or up-to-date this PDF is, but it should at least give you some idea as to what a SoC can do without bogging down the CPU cores: https://dl.radxa.com/rockpis/docs/hw/datasheets/Rockchip%20R... Check chapter 9 for example, all of those boxes are things you don't have to spend CPU cycles on. If you did use the CPU, it would be super slow.
For a hobby project I'm trying to find a solution - Power budget for multiple is sub 200w. Need to run inference on a lower resolution video stream (or multiple, that be nice) to do object detection. Cost is a factor because I need to have multiple angles to determine where in relation to a mobile platform the target object is. I'm looking at the Coral.ai board because RPi like boards lack the ability to do ML tasks at reasonable FPS and NVidia seems to have abandoned the lower cost side of the market since the Jetson Nanos seem to be less and less available. (Not that Coral.ai boards are available at all...)
Going to have to dig into the sensors they use - had passable luck with non ML tasks using dirt cheap camera modules from laptops running at low resolutions right up until I started moving the cameras at all and then it became a blurry mess because they were so small their exposure times were high. (I'm trying to also avoid having to put a bunch of illumination near the cameras so it doesn't entirely look like a biblically accurate angel)
EDIT:
* https://docs.luxonis.com/projects/api/en/latest/samples/Colo...
They have a home grown AI accelerator along with free Deep learning SDK.
Also offer a pretty easy online tool (free again) to use called TI Edge AI studio. They are using extending the existing AI solutions that come from the higher performance parts like TDA4 and AM68A parts. Pretty good considering a lot of these manufactures are just buying other AI Ip that isn’t performing great and investing in their own engineering.
nVidia's platform is just a huge mess. I tried to get their SDK running and their own documentation was out-of-date, missing necessary links, and sometimes blatantly wrong. I wasn't going to dump $400+ into that ecosystem.
Google gives up on hardware consistently and has the worst support of any existing software company (effectively zero) and has bungled the AI hand every change they get.
ARM NPUs I am not going to bother with. I can't even get video encoder acceleration working on a non-Pi ARM SoC except for the Rock64 and that is like 6 years old and was missing that functionality for 4 of them.
Intel only cares about its corporate partners and doesn't give a crap about hobbyists in regards to A.I. But their VPU was (is) decent and Oak guaranteed supply for at least until 2025 or thereabouts and built a useable API so we don't have to mess with OpenVINO.
It's all a mess right now but I can't say that competition is bad. It will be nice if we dispense with all the bespoke platforms and agree on some common architecture for edge devices, but I won't hold my breath.
So if you have a thing that has all the data and all the work done on an engine elsewhere and you just need to have a SoC to turn the thing on and off and get the data in and out, that's where a potato-SoC could work. Of course, the potato would need good distro support with up-to-date kernel, drivers, python, libraries etc. and if you're connecting it to the network, best make sure it's also getting patched consistently.
So, as far as ML Potatoes go, that's about it. It is a totally valid question by the way, even if asked in jest.
It's the same situation as all the bitcoin miners from years ago. They mined on the GPU, so the CPU didn't matter as much as how many PCIe slots the motherboard had. The CPU is purely there to drive the operating system and enable the network to communicate with the GPU.
If you aren't using the CPU for anything other than the PCIe slot to enable a Nvidia card for ML, then a very cheap potato makes a lot of sense.