Gemma 4 on Cerebras - The Fastest Inference Is Now Multimodal
cerebras.ai
cerebras.ai
1. a smaller model
2. also non local, hosted on cloud
I can't think of any case.
Text autocorrect on my phone? Like give it all the context about me and so on.
its not necessarily specifically labout gemma 4, but in a year or 2 when we have opus class models at 2000 tps imagine the productivity.
But for 2, probably only useful if you have a huge batch workload you want to get done quicker and don't want the local hardware for it?