If you look at A100 compared to V100 for e.g. FP32 FMA performance (not tensor). 14.1TFLOPS -> 19.5 = +38%, for 2x transistors (16->7nm), +35% SMs and 250W->400W is not that great. Note that NVIDIA uses boost clock for all A100 numbers and seems not have published any base clock so far. So their is a chance that actual sustained A100 performance is lower.
Turing GPUs have rather large dies. TU102 754nm vs GP102 471nm. So comparing them as is, isn't quite fair.
On the CPU front Intel used to use rather small dies for consumers (and even use die shrinks to just cram more chips onto a waver -> more $$$), but now that AMD forces their hand, they are giving in. But of course a lot this area goes into extra cores, not single threaded performance (diminishing returns there).