If you want to compare the host too, then you should also take into account the TPU host.
Each TPUv5p host has 208 vCPU, 448 GB ram, 200 Gbps NIC (for data loading), 4 TPUv5p chips (8 cores). Each TPU chip has 95GB HBM with 2765GBps memory bandwidth, 3D interconnect with 4800 Gbps (3D interconnect means every chip has 6x 800 Gbps feeds, each going to a different chip in the torus).
One big difference between TPUs and GPUs is that the former have few beefy cores (2x chips) while the latter a lot of smaller cores. It can help or hinder depending on the workload. It makes it simpler to do deterministic training, which helps both debugging and optimization.
So, each system has 448GB host RAM and 4 TPU chips with 380GB HBM. The amount of host memory doesn't really matter that much: it's there to make sure that accelerators can be fed data at full speed.
TPU5p pods are organized in "cubes" of 16 machines (64 TPUs) that can be assembled in any shape for different kinds of data/model parallelism (see twisted tori in the TPUv4 paper). Each cube has 6080GB (6TB) of HBM ram and 29.3 petaflops (bfloat16) of theoretical peak performance. You can get multiple exaflops from a single pod and use multiple pods in the same datacenter with multislice if needed.
TPU are more power efficient than GPUs (e.g. A100 has a TDP 3x of TPUv4 + 3D torus requires fewer connections that a fattree interconnect, and optics do consume their fair bit of power), so you can scale to higher compute per datacenter before you hit that bottleneck on really really large models.
Also notice that you are using bare metal numbers for Grace hopper and I can only quote (what is publicly available) userspace/VM numbers for TPUs. For example, the actual physical RAM on the hosts is higher than 448 GB; so 448 vs 480 is apples to oranges.
Source: https://cloud.google.com/tpu/docs/v5p-training + I am oncall for these platforms.
https://arxiv.org/ftp/arxiv/papers/2304/2304.01433.pdf for TPUv4 vs. A100
---
I believe Grace Hopper is an amazing platform. It is a more general one than TPUv5p pods. Each optimized for different constraints and TCO/perf.