GPU's Rival? What Is Language Processing Unit (LPU)
turingpost.com
turingpost.com
That being said, yes the software is impressive and the engineering team at Groq is top-notch.
Unfortunately, the chips are mainly aimed at inference, and it seems like a lot of the large investments at the moment are being driven by training.
EDIT: In the spirit of full disclosure, I suppose I should point out that I own lots of Groq shares, so have every interest in their success.
The compiler obviously knows how to schedule all this to produce good timing. That's why the utilization is so high. Every piece of the chip is being used. Whereas with a GPU model, you have to kind of play 'code golf' and have a deep understanding of the architecture and the execution engine to be able to determine when a memory access is going to cause a pipeline stall.
If you understand the ISA of the GroqChip, you understand completely how the chip works. There is no pipeline stall. There is no waiting for memory. There is no memory hierarchy. Even the interconnects work at the same speed as the chip (or something like that... it's all in the paper, so public information), so when they network chips together, the latency between any two subunits of every chip is already known by the network topology and the chip timings.
TL;DR It's all very deterministic and this works well for tensor operations. Parallelism is determined at compile time
Source: https://www.youtube.com/watch?v=pb0PYhLk9r8 . This is basically exactly how the chip works. It's not some simplified architectural or marketing diagram. It's literally how it works.
I mean this is a bit of a hyperbole. Even if things are deterministic, the compiler might not be able to schedule things optimally, their can be pipeline interlocks or memory bank conflicts (depending on sizes involved) in some scenarios.
1) AMD, Intel, or ARM buy Groq.
2) They become like ARM and license the ISA to system integrators.
3) They license a GPU core to integrate to their chips.
What other outcomes are likely?
For determining what goes into your faceplant feed it's perfect, for controlling your antilock brakes, not so much...
A lot of serious processing has nebulous results, like analyzing MRI images for signs of cancer, but not all computing falls into this category. Some things need to be deterministic.