I just tried it on their website and it is extremely fast. I wonder what is the value prop of this? Where would I want
1. a smaller model
2. also non local, hosted on cloud
I can't think of any case.
1. a smaller model
2. also non local, hosted on cloud
I can't think of any case.
But for 2, probably only useful if you have a huge batch workload you want to get done quicker and don't want the local hardware for it?
its not necessarily specifically labout gemma 4, but in a year or 2 when we have opus class models at 2000 tps imagine the productivity.
Text autocorrect on my phone? Like give it all the context about me and so on.