No HBM because they use tons of fast SRAM instead. Isn't that the main driver for performance here?
(the way I understood it => it's still cost effective at scale due to throughput increase this brings)
(the way I understood it => it's still cost effective at scale due to throughput increase this brings)
No doubt fast SRAM helps, but from a computation pov imho its that they've statically planned computation and eliminated all locks.
Short explainer here: https://www.youtube.com/watch?v=H77tV1KcWIE (Based on their paper).
Most important, even ignoring latency, is throughput (tokens) per $$$. And according to their own benchmark [1] (famous last words :)), they're quite cost efficient.
[1] https://www.semianalysis.com/p/groq-inference-tokenomics-spe...