The only non-TPU fast models I'm aware of are things running on Cerebras can be much faster because of their CPUs, and Grok has a super fast mode, but they have a cheat code of ignoring guardrails and making up their own world knowledge.
Where are you getting that? All the citations I've seen say the opposite, eg:
> Inference Workloads: NVIDIA GPUs typically offer lower latency for real-time inference tasks, particularly when leveraging features like NVIDIA's TensorRT for optimized model deployment. TPUs may introduce higher latency in dynamic or low-batch-size inference due to their batch-oriented design.
https://massedcompute.com/faq-answers/
> The only non-TPU fast models I'm aware of are things running on Cerebras can be much faster because of their CPUs, and Grok has a super fast mode, but they have a cheat code of ignoring guardrails and making up their own world knowledge.
Both Cerebras and Grok have custom AI-processing hardware (not CPUs).
The knowledge grounding thing seems unrelated to the hardware, unless you mean something I'm missing.
The citation link you provided takes me to a sales form, not an FAQ, so I can't see any further detail there.
> Both Cerebras and Grok have custom AI-processing hardware (not CPUs).
I'm aware of Cerebras' custom hardware. I agree with the other commenter here that I haven't heard of Grok having any. My point about knowledge grounding was simply that Grok may be achieving its latency with guardrail/knowledge/safety trade-offs instead of custom hardware.
I don't see any latency comparisons in the link
https://jax-ml.github.io/scaling-book/gpus/#gpus-vs-tpus-at-...
Re: Groq, that's a good point, I had forgotten about them. You're right they too are doing a TPU-style systolic array processor for lower latency.
For each you can use it as “instant” supposedly without thinking (though these are all exclusively reasoning models) or specify a reasoning amount (low, medium, high, and now xhigh - though if you do g specify it defaults to none) OR you can use the -chat version which is also “no thinking” but in practice performs markedly differently from the regular version with thinking off (not more or less intelligent but has a different style and answering method).
Coming up with all that fluff would keep my brain busy, meaning there's actually no additional breathing room for thinking about an answer.
> Coming up with all that fluff would keep my brain busy, meaning there's actually no additional breathing room for thinking about an answer.
It gets a lot easier with practice: your brain caches a few of the typical fluff routines.
And now with RAM, GPU and boards being a PitA to get based on supply and pricing - double middle finger to all the big tech this holiday season!
It's a lost battle. It'll always be cheaper to use an open source model hosted by others like together/fireworks/deepinfra/etc.
I've been maining Mistral lately for low latency stuff and the price-quality is hard to beat.
They do have a priority tier at double the cost, but haven't seen any benchmarks on how much faster that actually is.
The flex tier was an underrated feature in GPT5, batch pricing with a regular API call. GPT5.1 using flex priority is an amazing price/intelligence tradeoff for non-latency sensitive applications, without needing to extra plumbing of most batch APIs
Turns out becoming a $4 trillion company first with ads (Google), then owning everybody on the AI-front could be the winning strategy.