How to they run programs on it?
With a PC+GPU system you can just run the Python + tensorflow/pytorch/etc. code. All required drivers can be easily installed.
How to they do it with a custom co-processor?
How to they run programs on it?
With a PC+GPU system you can just run the Python + tensorflow/pytorch/etc. code. All required drivers can be easily installed.
How to they do it with a custom co-processor?
Once you have the two machines able to talk to one another on the physical level, you need to decide what kind of protocol they're going to use.
The simplest example of that whole flow is serial ports, which you may be familiar with if you've ever worked with arduino or microcontrollers or other embedded systems. So basically the answer is "the same way any two computers talk to one another", a protocol carried over some kind of physical medium. At the point you have to create your own drivers, and your own way of compiling code to run on the co-processor, but fundamentally it's no different from loading an arduino sketch onto a micro-controller.
There are also solutions involving shared memory, where all the processors/co-processors are connected to the same ram chips. Those generally aren't practical with x86 chips, although it's a popular way for baseboard processors to talk to the ARM chips that run cellphones. It's also ultimately they way "symmetric multiprocessing" works in every computer that has more than one core on it's CPU. It's also probably how the "neural processor" CPU communicates with all the custom silicone that actually enables this product. It's likely they have a conventional microcontroller core that handles the input stream from the host computer and puts that data into the RAM layout required by the "NPU". What they've labeled the "front end". I haven't really looked into their architecture in any depth though.
"Exotic" connections go beyond PCIe. You could theoretically share DDR4 RAM. With this, its probably more natural to have "cache-coherent" interconnects. Assuming the MESI protocol (aka: very much simplified), the coprocessor can raise a section of RAM (64 bytes or so) into the "exclusive" state. As long as the coprocessor is holding exclusive, the CPU will stall as it tries to read that RAM. The co-processor can return the memory to the CPU control by setting it to "Invalid" state. The CPU can then read (or write) to memory by setting the memory to exclusive: under the CPU's control.
Communication to and from the co-processor in this case isn't "memory-mapped IO", as much as it is "I/O over memory". :-). My explanation above is simplified, but hopefully you get the gist.
A realistic example of a coherency-fabric is something like POWER9 OpenCAPI can communicate with NVidia GPUs at over 50GBps, while sharing the same memory space. (PCIe 3.0 and PCIe 4.0 has features to allow a similar effect at slightly slower speeds. I would bet that PCIe 5.0 is working on some kind of coherency fabric, but I don't actually know).
The most exotic would probably be iGPUs (Intel) and APUs (AMD), which perform this kind of cache-coherent communication but at the L3 cache level and below. Since APUs and iGPUs share the same chip, they can send these messages back and forth without ever leaving the chip. Just entirely within L3 cache itself.
RAM has always forced CPUs to stall: refresh events are the most common (RAM is refreshing the voltage on all memory: it takes a long time and the CPU may have to wait in rare cases). But modern multicore CPUs have to have this coherency fabric to ensure that the different cores see memory in the correct order.
It doesn't really matter if the CPU is stalling because of a memory refresh, or a CPU-exclusive hold on some RAM... or even I/O pretending to be a CPU-core holding some RAM. The CPU's method of attack remains the same.
-------
This is most clear with Intel's Optane DIMMs, persistent storage that operates over the DDR4 protocol, pretending to be RAM. Slower than true DDR4 yes, but fast enough that it wanted to move off of PCIe and pretend to be DDR4 RAM. I expect more and more I/O to be "pretending to be RAM" in some effect, either through a coherency fabric or maybe even just copying the DDR4 protocol (like Optane).
> The ISA consists of instructions with up to 4 slots with complex values. There are only eight instructions in total – two DMA read and write, three dot-product operations, and scaling and element-wise addition.
That's not something you're going to target with a standard compiler. It doesn't even have any control flow instructions. Effectively they've built a "convolve machine": all they have to do is set up the addresses for the areas of memory containing the image and the weights, let it run, and collect the result.
But I might be wrong...