TPU v4 provides exaFLOPS-scale ML with efficiency gains
cloud.google.com
cloud.google.com
That said, it's also incredible how fast things move in this space:
> Midjourney, one of the leading text-to-image AI startups, have been using Cloud TPU v4 to train their state-of-the-art model, coincidentally also called “version four”.
Midjourney is already on v5 as of the date of publication of this press release.
JAX will compile your model with XLA today, letting you run it on Nvidia, Apple Silicon (CPU for now), google's TPU, etc.
Overall, XLA really shines on TPUs, is doing very well on GPUs and sadly lags behind on CPUs (reaching about single core performance).
As a bonus, JAX is basically Numpy++: you can use it to build scientific/numerical software and run it easily on GPU and TPU!
The lock-in is bad enough when dealing with niche hardware on-prem, i certainly won't deal with niche hardware in the cloud.
(With Altra availability ramping up)
https://www.ipi.wiki/products/com-hpc-ampere-altra?variant=4...
This sounds quite bad in a press release when Midjourney is at v5. Why did they move away?
> The Google tradition is to write retrospective papers after systems are deployed and running production apps. Both TPU v4s and A100s deployed in 2020 and both use 7nm technology. The newer, 700W H100 was not available at AWS, Azure, or Google Cloud in 2022. The appropriate H100 match would be a successor to TPU v4 deployed in a similar time frame and technology (e.g., in 2023 and 4 nm).
However, I remain a bit skeptical of the business case for TPUs for 3 core reasons:
1) 100000x lower unit production volume than GPUs means higher unit costs
2) Slow iteration cycle - these TPUv4 were launched in 2020. Maybe Google publishes one gen behind, but that would still be a 2-3 year iteration cycle from v3 to v4.
3) Constant multiple advantage over GPUs - maybe 5-10x compute advantage over off the shelf GPU, and that number isn't increasing with each generation.
It's cool to get that 5-10x performance over GPUs, but that's 4.5yrs of Moore's Law, and might already be offset today due to unit cost advantages.
If the TPU architecture did something to allow fundamentally faster transistor density scaling, it's advantage over GPUs would increase each year and become unbeatable. But based on the TPUv3 to TPUv4 perf improvement over 3 years, it doesn't seem so.
Apple's competing approach seems a bit more promising from a business perspective. The M1 unifies memory reducing the time commitment required to move data and switch between CPU and GPU processing. This allows advances in GPUs to continue scaling independently, while decreasing the user experience cost of using GPUs.
Apple's version also seems to scale from 8GB RAM to 128GB meaning the same fundamental process can be used at high volume, achieving a low unit cost.
Are there other interesting hardware for ML approaches out there?
Google also has Coral, which is a non-cloud mobile-focused TPU that you can buy and plug in (USB or PCIe).
Naturally, this is orders of magnitude less powerful than the kind of TPU being discussed here but I just like the idea of local compute.
It is completely unreasonable to expect something like that.
This is obviously an exaggeration, I wonder what the actual ratio is between TPU proudction and e.g. A100 production.
Two points. Nvidia's RTX 3000 series (3060 Ti, 3080, and many other flavors) ships 6 or more flavors per generation. The related silicon has names like the GA102, GA103, GA104, GA106, and GA107. So only 1/6th of the consumer market for Nvidia silicon can be amortized over any single design.
I wouldn't be at all surprised to see Google making the TPUs by the million. I found a vague reference to 9 exaflops and single facilities (one of many) costing $4 billion to $8 billion.
So I wouldn't assume that the consumer GPU market/number of silicon designs is 100,000 times larger than the TPUv4 market.
> Slow iteration cycle
True. Then again generations make much less difference than they used to. Gone are the days where even after a multiple generations that average performance increases by 2x. Sure nvidia's 4000 series claims 2x ... on raytracing. But normal game performance seems to be more like 15%. Sure various trickery like DLSS helps, but similar tricks are increasing the performance of older cards as well. Similarly apple's a14 -> a15 -> a16 (or m1 -> m2 if you prefer) chips have had modest performance increases and mostly have increases in perf/watt.
> 4.5yrs of Moore's Law
It's dead Jim.
The picture in this article shows 8 racks with (according to the paper on arxiv) has 16 TPU sleds each, 4 TPUs per sled. that's only 512 chips. According to the paper, it is one of eight in a 4096-chip supercomputer. If you give them 10-100 of those around the world, you get 40,000-400,000 chips. That's enough for reasonable scale. Nvidia should still have 100x (or more) their scale.
The advantage of this hardware isn’t just the raw compute capacity. It’s the massive IO bandwidth.
GPUs are fast, but they’re not designed for massive all-to-all communication across large networks. That’s where modules like this shine.
Is that substantially different from the Nvidia release cadence? According to the Wikipedia announcement dates, their last three generations were H100 in March 2022, A100 in May 2020, V100 in March 2017.
This blog [1] does make the argument that M1 Max has about 8x lower performance than a consumer grade Nvidia GPU 2080, but also uses 8x less watts, so it's possible Apple can make a product with similar performance/watt. However, I would say that they are farther away than TPUv4 for now, because the product doesn't exist.