Wow. I knew custom Google silicon was used for inference, but I didn't realize it was used for training too. Does this mean Google is free of dependence on Nvidia GPUs? That would be a huge advantage over AI competitors.
Wow. I knew custom Google silicon was used for inference, but I didn't realize it was used for training too. Does this mean Google is free of dependence on Nvidia GPUs? That would be a huge advantage over AI competitors.
(TPUs have had BarnaCore for efficient embedding lookups since TPU v3)
BarnaCore existed and was used, but was tailored mostly for embeddings. BTW, IIRC they were called that because they were added "like a barnacle hanging off the side".
The evolution of TPU has been interesting to watch; I came from the HPC and supercomputing space, and seeing Google as mostly-CPU for the longest time, and then finally learning how to build "supercomputers" over a decade+ (gradually adding many features that classical supercomputers have long had), was a very interesting process. Some very expensive mistakes along the way. But now they've paid down almost all the expensive up-front costs and can now ride on the margins, adding new bits and pieces while increasing the clocks and capacities on a cadence.
The superscalers are all working on this. https://aws.amazon.com/ai/machine-learning/trainium/
(PS: we are startup trying to make TPUs more accessible, if you wanna fine-tune Llama3 on TPU check out https://github.com/felafax/felafax)
Of course this does not mean their models will necessarily be proportionally better, nor does it mean Google won't buy GPUs for other reasons (like providing them to customers on Google Cloud.)
Nvidia is making big bucks "selling shovels in a gold rush". Google has made their own shovel factory and they can avoid paying Nvidia's margins.
+1 on this. The tooling to use TPUs still needs more work. But we are betting on building this tooling and unlocking these ASIC chips (https://github.com/felafax/felafax).
https://cloud.google.com/tpu/docs/v6e
The A100 (from 2020) had 300 TFLOPs (1.5x) and 80GB HBM. It's only now with Trillium that they are starting to actually beat Nvidia in a chip vs chip battle, but being faster per chip was never the point. TPUs were supposed to be mass produced and connected into pods so that they become cheaper and they indeed are.
Source? This would make quite the splash in the market
Llama 2 was released well over a year ago and was training between Meta and Microsoft.
They could train something similar, but it'd be super weird if they called it Llama 2. They could call it something like "Gemini", or if it's open weights, "Gemma".