Sure, I think having dedicated acceleration blocks for transformers explains that pretty well, and is definitely exciting for people who can take advantage of it.
I was more contextualizing this, where a potential 75% increase in power (on a smaller node) to deliver 50% more performance is less impressive:
> I think it's even more impressive that the smallest increase vs. A100 is still 50% (for ResNet)
If I were deploying 8-GPU nodes today, I'm not sure whether I would want to pick between 4-GPU nodes or doubling my node TDP. The DGX A100 and H100 both have 8 cards, and 640GB of VRAM, but peak power increased from 6.2kW -> 10.2kW. If you're training giant models and constrained on VRAM, you'll potentially need just as many nodes as before just to fit your models into memory. Your training will be faster but it's hard to avoid the power increase.
The power increase is much more reasonable if you stick with PCIe SKUs (300W -> 350W, literally half the power draw of the SXM SKU), but you pay for that with 20% less compute and 33% less memory bandwidth.