ASIC: Application-Specific Integrated Circuit: https://en.wikipedia.org/wiki/Application-specific_integrate...
FPGA: Field-Programmable Gate Array: https://en.wikipedia.org/wiki/Field-programmable_gate_array
Graphcore IPU: https://en.wikipedia.org/wiki/Graphcore
"Intel’s Exascale Dataflow Engine Drops x86 and von Neumann" (2018) https://www.nextplatform.com/2018/08/30/intels-exascale-data... :
> The new architecture that they have dreamed up is called a Configurable Spatial Accelerator, but the word accelerator is a misnomer because what they have come up with is a way to build either processors or coprocessors that are really dataflow engines, not serial processors or vector coprocessors, that can work directly on the graphs that programs create before they are compiled down to CPUs in a traditional sense. The CSA approach is a bit like having a switch ASIC cross-pollenate with a math coprocessor, perhaps with an optional X86 coprocessor perhaps in the mix if it was needed for legacy support.
> The concept of a dataflow engine is not new. Modern switch chips work in this manner, which is what makes them programmable to a certain extent rather than just static devices that move packets around at high speed. Many ideas that Intel has brought together in the CSA are embodied in Graphcore’s Intelligence Processing Unit, or IPU. The important thing here is that Intel is the one killing off the X86 architecture for all but the basic control of data flow and moving away from a strict von Neumann architecture for a big portion of the compute. So perhaps Configurable Spatial Architecture might have been a better name for this new computing approach, and perhaps a Xeon chip will be thought of as its coprocessor and not the other way around.
FWIU the Intel 486 / 487 was the first CPU with a math coprocessor.
Coprocessor: https://en.wikipedia.org/wiki/Coprocessor
With some newer architectures, the GPU(s) are directly connected to HBMe RAM; which somewhat eliminates the CPU and Bus performance bottlenecks that ASICs and FPGAs are used to accelerate beyond.
High Bandwidth Memory > Technology > HBM3E: https://en.wikipedia.org/wiki/High_Bandwidth_Memory#Technolo...
AI Accelerator: https://en.wikipedia.org/wiki/AI_accelerator
Dataflow architecture: https://en.wikipedia.org/wiki/Dataflow_architecture
Massively parallel processor array: https://en.wikipedia.org/wiki/Massively_parallel_processor_a... :
> MPPAs are used in high-performance embedded systems and hardware acceleration of [various workloads], which otherwise would use FPGA, DSP and/or ASIC chips.
> These processors pass work to one another through a reconfigurable interconnect of channels.
"A PCIe Coral TPU Finally Works on Raspberry Pi 5" https://news.ycombinator.com/item?id=38310063 :
> An HBM3E HAT would or would not yet make TPUs more useful with a Raspberry Pi 5?