What's interesting about this demo is the speed at which it is running, which demonstrates the "Groq LPU™ Inference Engine".
That's explained here: https://groq.com/lpu-inference-engine/
> This is the world’s first Language Processing Unit™ Inference Engine, purpose-built for inference performance and precision. How performant? Today, we are running Llama-2 70B at over 300 tokens per second per user.
I think the LPU is a custom hardware chip, though the page talking about it doesn't make that as clear as it could.
https://groq.com/products/ makes it a bit more clear - there's a custom chip, "GroqChip™ Processor".