They get impressive performance, but it's not really the first "LLM-native processor". They built the thing a while ago before LLMs were that hot, and rebranded it as an LPU to pivot into Generative AI: https://www.eetimes.com/groq-demos-fast-llms-on-4-year-old-s...
Aren't TPU the first LLM-Native processors?
TPUs are for matrix multiplication and nn in general. They have a model compiler so stuff runs on their hardware. LLMs are just the new hotness so they’re the current focus.
This article has a strong tone of doubt that it's real- which is a bit odd. Am I missing something? This isn't a magic leap situation- they demonstrated it working, right?
They demonstrated that throwing $12 million worth of silicon at the problem beats a benchmark, which, of course, it does.
It feels like it's pre-Cryptocurrency mining boom.
Most of the people tried to get the proprietary GPU, and somebody makes an FPGA & ASIC for Accelerated devices.
Hope this company’s IP security is like Fort Knox because you know China is already trying to get access to everything they got.
I foresee this being the game changer that will see more companies selfhosting their own AI applications. If you can serve an entire office with just one or two LPUs, you're kinda stupid to not do it.
You need an entire rack or two worth of Groq systems to get this performance, because the only memory the system has is on-chip SRAM (220MB). If you're trying to run on a small number of systems, you need local DRAM or HBM, which Groq doesn't have.