The book is definitely not closed, but the other limits are somehow less problematic than the license-based downclocking.
You you use more power (and get hotter temps: these are exactly proportional, so you can mostly just talk about them as one) with wider vectors because you are doing more work. When you look at it on a per-element basis, you use less power per element with wider vectors. E.g., you might use 1 pJ per element for 256-bit FMA but only 0.8 for 512-bit FMA.
Of course, since you can do 2x as many total elements in 512-bits on a 2 FMA machine, you can be both more efficient but use more total power, so you can get TDP or thermal limits with 512-bit code that you wouldn't on 256-bit, but it should still per faster and more efficient per element.
All of this assumes you can usefully use the 2x more work with the larger vectors. Sometimes the scaling is worse: e.g., for lots of short arrays, when a lookup table is involved, when additional shuffling or transposition is required with larger vectors, etc. In that case you could end up less efficient with larger vectors.