If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone.
From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific GPU's.
If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone.
From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific GPU's.
[1] https://mlcommons.org/benchmarks/training/
[2] https://cloud.google.com/blog/products/compute/introducing-t...
jonathan [at] tensordock.com
It should be possible to hook up your idle devices to Nautilus (https://nationalresearchplatform.org/), which is a cluster set up to support researchers at a bunch of universities. I can't guarantee anything since I'm not involved in the cluster management itself, but I can put you in contact with those who are if you're interested.
Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.
I would bet money that TPUs are at least better at doing AI research than anything Nvidia will sell you. That alone might be enough for Google to keep getting some new ones fabbed each year. The TPUs you can rent on Google Cloud might very well just be hardware requisitioned by the AI team, for the AI team, that they aren't always using to capacity, and so is "earning out" its CapEx through public rentals.
TPUs are maybe also better at other things Google does internally, too. Running inference on YouTube's audio+video-input timecoded-captions-output model, say.
"Deployed since 2020, TPU v4 outperforms TPU v3 by 2.1x and improves performance/Watt by 2.7x. ... For similar sized systems, it is ~4.3x--4.5x faster than the Graphcore IPU Bow and is 1.2x--1.7x faster and uses 1.3x--1.9x less power than the Nvidia A100. TPU v4s inside the energy-optimized warehouse scale computers of Google Cloud use ~2--6x less energy and produce ~20x less CO2e than contemporary DSAs in typical on-premise data centers."
Here is a link to the paper: https://dl.acm.org/doi/pdf/10.1145/3579371.3589350
Which sure makes the H100 sound both faster and more efficient (per unit of compute) than the TPU v4, given what was in your quote. I don't think your quote does anything to support the position that TPUs are noticeably better than Nvidia's offerings for this task.
Complicating this is that the TPU v5 generation has already come out, and the Nvidia B100 generation is imminent within a couple of months. (So, no, a comparison of TPUv5 to H100 isn't for a future paper... that future paper should be comparing TPUv5 to B100, not H100.)
[0]: https://developer.nvidia.com/blog/nvidia-hopper-architecture...
Despite various details I don't think that this is an area where Facebook is very different from Google. Both have terrifying amounts of datacenter to play with. Both have long experience making reliable products out of unreliable subsystems. Both have innovative orchestration and storage stacks. Meta hasn't published much or anything about things like reconfigurable optical switches, but that doesn't mean they don't have such a thing.
I think it's aimed at medium to large corporations.
Massive corporations such as Meta and OpenAI would build their own cloud and not rely on this.
The GPU really is a shovel, and can be used without any subscription.
Don't get me wrong, I want there to be competition with Nvidia, I want more access for open source and small players to run and train AI on competitive hardware at our own sites.
But no one is competing, no one has any idea what they're doing. Nvidia has no competition whatsoever, no one is even close.
This lets Nvidia get away with adding more vram onto an AI specific GPU and increase the price by 10x.
This lets Nvidia remove NVLink from current gen consumer cards like the 4090.
This lets Nvidia use their driver licence to prevent cloud platforms from offering consumers cards as a choice in datacenters.
If Nvidia had a shred of competition things would be much better.
While I do not actually think Google's chips are better or close to being better, I don't think this actually holds?
If the upside of <better chip> is effectively unbounded, it would outweigh the short term benefit of selling them to others, I would think. At least for a company like Google.
They do sell them - but through their struggling cloud business. Either way, Nvidia's margin is google's opportunity to lower costs.
> I can see Google's custom chips are 15x to 30x slower to train AI
TPUs are designed for inference not training - they're betting that they can serve models to the world at a lower cost structure than their competition. The compute required for inference to serve their billions of customers is far greater than training costs for models - even LLMs. They've been running model inference as a part of production traffic for years.
(I work for Google, but the above is public information.)
Nvidia is in a vaguely unique position in that their products have great tooling support and few companies sell silicon at their scale.
In AI, Google (TPU) and Intel (Gaudi) each have chips they push in cloud offerings. The cloud offerings have cross selling opportunities. That by itself would be a reason to keep it internal at their scale. It might also be easier to support one, or a small set, of deployments that are internal vs the variety that external customers would use.