TPU transformation: A look back at 10 years of our AI-specialized chips
cloud.google.com
cloud.google.com
TPUs are the second most widely used environment for training after Nvidia. It's the only environment that people build optimized kernels for outside CUDA.
If it was separate to Google then there a bunch of companies who would happily spend some money on a real, working NVidia alternative.
It might be profitable from day one, and it surely would gain substantial market capitalization - Alphabet shareholders should be agitating for this!
LLM AI is largely HBM bottlenecked anyway i.e. Samsung, SK Hynix and Micron are where the supply chain limits enter the picture.
The memory industry just got busted from the covid bubble and are not too keen to jump into the AI bubble.
Edit: But perhaps with the exclusivity deals, the likes of TSMC are less reliant on spreading the cost over 15+ years than they used to be. To be clear, I was talking about long-term use.
But it looks like they've actually been buying back some shares - they've got fewer shares outstanding than they did a year or two ago.
Not that it matters much - they've still got plenty of cash and other capital available.
We don't even have enough McDonald's employees, how the hell are we going to just suddenly have multiple companies creating fabs left and right? TSMC cannot even build their Arizona plant without a shortage of workers.
There is good reason nobody wants to be in the fab business.
Between that and the fact Google already sells "Coral Edge TPUs" [1] I'd think they could manage to untangle things.
Whether the employees would want to be spun off or not is a different matter, of course...
For a large, established, quasi-monopoly company it's always more attractive to keep things inside their walled gardens. Suggesting that Google should start supporting TPUs outside Google Cloud is like suggesting that Apple should start supporting iOS on non-Apple hardware.
I think nvidia is ecstatic about having commoditised their complement, and having the only ML acceleration option that's available from every cloud provider and on-prem.
Why have Amazon, Google and Microsoft as competitors when you can have them as customers instead?
Not really. Google TPUs require google's specific infrastructure, and cannot be deployed out side the Google Datacenter. The software is google specific, the monetization model is google specific.
We also have no idea how profitable TPUs would actually be if a separate company. The only customer of TPUs is Google and Google Cloud.
Nvidia has spent 20 years on this which is why they're good at it.
> If it was separate to Google then there a bunch of companies who would happily spend some money on a real, working NVidia alternative.
Unfortunately, most people really don't care about Nvidia alternatives, actually -- they care about price, above all else. People will say they want Nvidia alternatives and support them, then go back to buying Nvidia the moment the price goes down. Which is fine, to be clear, but this is not the outcome people often allude to.
Google could spin the TPU division out of Google, but 99% of the time people refer to moves like that they omit the implied follow up sentence which is "I can then buy a TPU with my credit card off the shelf from a website that uses Stripe." It is just not that simple or easy.
Not saying it is easy or to do it magically.
Just noting that Groq (founded by the TPU creator) did exactly this.
Building out your own vertically integrated offering with APIs is comparatively a lot simpler and significantly less risky in the grand scheme. For one thing, cloud APIs naturally benefit from the opex vs capex distinction that is often brought up here -- this is a big sales barrier, and thus a big risk. This is important because you can flush mid-8-figures down the toilet overnight for a single set of photomasks, so you are burning significant capital way before your foot is ever close to the proverbial door, much less inside it. You aren't going to make that money back selling single PCIe cards to enthusiastic nerds on Hacker News; you need big fish. Despite allusions to the contrary (people beating down your door to throw you bathtubs of money with no question), this isn't easy.
Another good example of verticality is the software. The difference in scope and scale between "Tools that we run" and "Tools you can run" is actually huge. Think about things like model choice -- it can be much easier to support things like new models when you are taking care of the whole pipeline and a complete offering, versus needing to support compiler and runtime tools that can compile arbitrary models for arbitrary setups. You can call it cutting corners, but there's a huge amount of tricky problems in this space and the time spent on procedural stuff ("I need to run your SDK on a 15 year old CentOS install!") is time not spent on the core product.
There are other architectural reasons for them to go this route that make sense. But I really need to stress here that a big and important one is that hardware is, in fact, a very difficult business even with a great product.
(Disclosure: I used to work at Groq back in 2022 before the Cloud Compute offering was available and LLMs were all the rage.)
I think some (large) buyers will want on-prem and they have large enough budgets to make that worthwhile.
I don't think "sell individual TPUs to random people" is a great model. Most are better served by the cloud rental approach (although they might not think so themselves).
AWS and Azure (to a lesser extent) can also make this argument.
Google is going to dominate LLM ushered AI era. Google has been AI first since 2016, they just don't have the opening. Sam, as inapt at engineering, just has no idea how to navigate the delicate biz & eng competitions.
See https://cloud.google.com/blog/products/compute/introducing-t... and https://cloud.google.com/blog/topics/systems/the-evolution-o...
It was OpenAI that showed you can actually deploy a large model, like GPT-4, to a large audience. Maybe Google didn't reach the cost efficiency with just internal use that NVIDIA does.
https://www.youtube.com/watch?v=nR74lBO5M3s
(note the lede on the TPU is buried pretty deep here)