Even then, I think that their primary use case is going to be consumer grade good AI on phones. I dunno why Gemma QAT model fly so low on the radar, but you can basically get full scale Llamma 3 like performance from a single 3090 now, at home.
Even then, I think that their primary use case is going to be consumer grade good AI on phones. I dunno why Gemma QAT model fly so low on the radar, but you can basically get full scale Llamma 3 like performance from a single 3090 now, at home.
Google has already started the process of letting companies self-host Gemini, even on NVidia Blackwell GPUs.
Although imho, they really should bundle it with their TPUs as a turnkey solution for those clients who haven't invested in large scale infra like DCs yet.
And also, Google's track record with hardware.
My point was a bit more specific though. To elaborate, I know of a number of publicly traded companies (USD $200M+ market cap) globally which have identified use cases for onprem AI and want to implement them actively but cannot, because they lack the knowhow to work with onprem, and hiring talent to implement that is just extremely expensive. Google should simply provide it as a turnkey bundle and milk them for it.