The thing about TPU clusters that they have hyper-torus optical interconnect between TPUs. This allows for extremely efficient weight updates. To replicate this with A100s you need very custom hardware/software deployment.
But to be fair, I don’t know what is latest and greatest available from NVidia or other clouds in this area right now.
EDIT: Looks like NVidia has NVSwitch, which provides interconnect for 256 GPUs. Pretty cool!