It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.
I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.
We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.