Trillium TPU Is GA
cloud.google.com
cloud.google.com
Wow. I knew custom Google silicon was used for inference, but I didn't realize it was used for training too. Does this mean Google is free of dependence on Nvidia GPUs? That would be a huge advantage over AI competitors.
Source? This would make quite the splash in the market
Llama 2 was released well over a year ago and was training between Meta and Microsoft.
They could train something similar, but it'd be super weird if they called it Llama 2. They could call it something like "Gemini", or if it's open weights, "Gemma".
(TPUs have had BarnaCore for efficient embedding lookups since TPU v3)
BarnaCore existed and was used, but was tailored mostly for embeddings. BTW, IIRC they were called that because they were added "like a barnacle hanging off the side".
The evolution of TPU has been interesting to watch; I came from the HPC and supercomputing space, and seeing Google as mostly-CPU for the longest time, and then finally learning how to build "supercomputers" over a decade+ (gradually adding many features that classical supercomputers have long had), was a very interesting process. Some very expensive mistakes along the way. But now they've paid down almost all the expensive up-front costs and can now ride on the margins, adding new bits and pieces while increasing the clocks and capacities on a cadence.
The superscalers are all working on this. https://aws.amazon.com/ai/machine-learning/trainium/
Of course this does not mean their models will necessarily be proportionally better, nor does it mean Google won't buy GPUs for other reasons (like providing them to customers on Google Cloud.)
Nvidia is making big bucks "selling shovels in a gold rush". Google has made their own shovel factory and they can avoid paying Nvidia's margins.
+1 on this. The tooling to use TPUs still needs more work. But we are betting on building this tooling and unlocking these ASIC chips (https://github.com/felafax/felafax).
https://cloud.google.com/tpu/docs/v6e
The A100 (from 2020) had 300 TFLOPs (1.5x) and 80GB HBM. It's only now with Trillium that they are starting to actually beat Nvidia in a chip vs chip battle, but being faster per chip was never the point. TPUs were supposed to be mass produced and connected into pods so that they become cheaper and they indeed are.
(PS: we are startup trying to make TPUs more accessible, if you wanna fine-tune Llama3 on TPU check out https://github.com/felafax/felafax)
If TPU's are really that good why on earth would google not sell them. People say its better to rent, but how can that be true when you look at the value of nvidia.
On the GPU side - One way to look at it is lets just ignore whether you're selling chips. Let's just look at where compute is done - because Google let's you rent time on the TPUs so if they're really great people will just choose to rent time on them. How much time is being rented on TPUs in the cloud vs. Nvidia GPUs? Google's TPUs are not a market leader, they're probably not even significant market share. Pretend Google decided to sell TPUs tomorrow - are many people who weren't choosing to rent TPUs going to decide to buy them? No. There's an element of interoperability and supply chain etc. etc. but I think the core of it is a full ML product the TPU isn't anywhere close to the Nvidia GPUs.
The only way the current situation makes sense is if your last sentence is true, which is what i'm getting at. Either TPUs are significantly behind Nvidia chips, or google is choosing not to add like a trillion dollars of market cap to their business its really one or the other.
Hence, buy $GOOG.
Right now if Google can earn more from using their TPUs than they would selling them (ie profitably utilize them) they would be crazy to sell them.
Companies are valued not on their present value, but the net present value of their future earnings. So the big multiple implies that there is much more revenue growth potential here. For example, NVidia has Cuda, Google doesn’t have anything as good for third party users. Maybe the market is pricing in that Google doesn’t have a good path to manufacture and sell these chips externally, while NVidia has room to grow into every 1GW datacenter that’s built over the next few decades.
Reminds me of what happened with Bitcoin mining ASICs. Everyone capable of making them realized they could make more money using them directly vs. selling them to others.
nVidia sells nearly 4M GPUs per year. Google claims like 100K TPUs. Scaling a production line by 10x is very difficult and Google has not shown aptitude in this area of expertise.
Even if Google wanted to scale 10x I'm not sure they could. nVidia is believed to be taking like half of TSMC's new capacity (existing capacity is not idle). I suppose technically that means the other half could be consumed by Google but it's likely TSMCs other customers wouldn't appreciate that.
Plus big tech companies have the data and customers and will probably be the only surviving big AI training companies. I doubt startups can survive this game - they can’t afford the chips, can’t build their own, don’t have existing products to leech data off of, and don’t have control over distribution channels like OS or app stores
Well, look at it this way. Nvidia played their cards so well that their competitors had to invent entirely new product categories to supplant their demand for Nvidia hardware. This new hardware isn't even reprising the role of CUDA, just the subset of tensor operations that are used for training and AI inference. If demand for training and inference wanes, these hardware investments will be almost entirely wasted.
Nvidia's core competencies - scaling hardware up and down, providing good software interfaces and selling direct to consumer are not really assailed at all. The big lesson Nvidia is giving to the industry is that you should invest in complex GPU architectures and write the software to support it. Currently the industry is trying it's hardest to reject that philosophy, and only time will tell if they're correct.
Interesting take, but why would demand for training and inference wade? This seems like a very contrarian take.
My overall point is that I think Nvidia played smartly from the start. They could derive profit from any sufficiently large niche their competitors were too afraid to exploit, and general purpose GPU compute was the perfect investment. With AMD, Apple and the rest of the industry focusing on simpler GPUs, Nvidia was given an empty soapbox to market CUDA with. The big question is whether demand for CUDA can be supplanted with application-specific accelerators.
At least for AI workloads, Google's XLA compiler and the JAX ML framework have reduced the need for something like CUDA.
There are two main ways to train ML models today:
1) Kernel-heavy approach: This is where frameworks like PyTorch are used, and developers write custom kernels (using Triton or CUDA) to speed up certain ops.
2) Compiler-heavy approach: This uses tools like XLA, which apply techniques like op fusion and compiler optimizations to automatically generate fast, low-level code.
NVIDIA's CUDA is a major strength in the first approach. But if the second approach gains more traction, NVIDIA’s advantage might not be as important.
And I think the second approach has a strong chance of succeeding, given that two massive companies—Google (TPUs) and Amazon (Trainium)—are heavily investing in it.
(PS: I'm also bit biased towards approach 2), we build llama3 fine-tuning on TPU https://github.com/felafax/felafax)
Nvidia also need to invent smth then, as pumping mining (or giving good to gamers) again is not sexy. What's next? Will we finally compute for drug development and achieve just as great results as with chatbots?
They do! Their research page is well worth checking out, they wrote a lot of the fundamental papers that people cite for machine learning today: https://research.nvidia.com/publications
> Will we finally compute for drug development and achieve just as great results as with chatbots?
Maybe - but they're not really analogous problem spaces. Fooling humans with text is easy - Markov chains have been doing it for decades. Automating the discovery of drugs and research of proteins is not quite so easy, rudimentary attempts like Folding@Home went on for years without any breakthrough discoveries. It's going to take a lot more research before we get to ChatGPT levels of success. But tools like CUDA certainly help with this by providing flexible compute that's easy to scale.
In drug discovery, we'd love to be able to show that virtual screening really worked- if you could do docking against a protein to find good leads affordably, and also ensure that the resulting leads were likely to pass FDA review (IE, effective and non-toxic), that could potentially greatly increase the rate of discovery.
Isn't it telling when Google's release of an "AI" chip doesn't include a single reference to nvidia or its products? They're releasing it for general availability, for people to build solutions on, so it's pretty weird that there isn't comparisons to H100s et al. All of their comparisons are to their own prior generations, which you do if you're the leader (e.g. Apple does it with their chips), but it's a notable gap when you're a contender.
Suspiciously there is a filter for "TPU-trillium" in the training results table, but no result using such an accelerator. Maybe results were there and later redacted, or have been embargoed.
The future of AI is local, not remote.
Would be neat if anyone has benchmarks!!
It's not obvious to me that a hardware business manufacturing tpus for the general public is necessarily more valuable than one that benefits from tight integration with Google's internal software stack and datacenter tech
It's orthogonal: tpu doesn't do much marketing and doesn't have to. Engineers at Google will use it because they have to and because of the huge cost advantage they get for using it. TPU is probably lacking the libraries that Nvidia has developed to be accessible to a broad array of use cases
The whole TPU line makes no sense to me, if its as good as it says (which is does seem to be) sell it publicly and add a trillion to your market cap.
The only way this makes any sense is if inside google people legitimately think google cloud is going to be bigger than nvidia+ a lot of azure&aws, which seems crazy
Google has vastly different constraints than an average startup or user of machine learning.
The flexibility of Nvidia GPU's, the software, these things are really valuable for most people. Only recently with LLMs has the majority of uses started to look very very similar (some slightly modified transformer)
Early versions of TPU required you to write your ml logic using a very restricted subset of tensorflow. (It had to be somewhat functional, etc)
Normal people don't write code like that, and it's not worth it for them to re write a working model to run on a TPU because software engineers are expensive and Nvidia GPUs are really good and general.
[1] Dataflow architecture:
https://en.wikipedia.org/wiki/Dataflow_architecture
[2] The GPU is not always faster:
... "The growing importance of multi-step reasoning at inference time necessitates accelerators that can efficiently handle the increased computational demands."
Unlike others, my main concern with AI is any savings we got from converting petroleum generating plants to wind/solar, it was blasted away by AI power consumption months or even years ago. Maybe Microsoft is on to something with the TMI revival.
Just that alone feels like an absolutely massive load to bear! But its only a drop in the bucket compared to everything else around this stuff.
But while I may be thirsty and hungry in the future, at least I will (maybe) be able to know how many rs are in "strawberry".
I don't see our current energy production scaling to meet the demands of AI. I see a lot of signals that most AI players feel the same. From where I'm sitting, AI is already accelerating energy generation to meet demand.
If your goal is to convert the planet to clean energy, AI seems like one of the most effective engines for doing that. It's going to drive the development of new technologies (like small modular nuclear reactors) pushing down the cost of construction and ownership. Strongly suspect that, in 50 years, the new energy tech that AI drives development of will have rendered most of our current energy infrastructure worthless.
We will abandon current forms of energy production not because they were "bad" but because they were rendered inefficient.