As it is, it seems the improvements are about making the hardware cheaper (as in capex, not opex).
This is just feels from me from what I hear on the news and see on the products though.
As it is, it seems the improvements are about making the hardware cheaper (as in capex, not opex).
This is just feels from me from what I hear on the news and see on the products though.
I think really good efficiency is possible right now, but the GPU makers don't want to make their consumer GPUs too good for AI - if the cards were more efficient, it'd be much easier to run multiple - while data center ones have a bunch of additional power overheads.
Especially when chips are becoming cheaper (in the capex sense), then you can afford to only run them when power is cheap.
Btw, from where do you take the notion that performance per Watt ain't increasing? We are also still using what's more or less general purpose GPU hardware; we could get a lot further if we were willing to specialise more. Which would be the natural avenue to explore, if progress in general purpose hardware slows down. Google is already looking.
GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs:
2020 RTX 3090: 350W
2022 RTX 4090: 450W
2025 RTX 5090: 575W
Bigger dies means lower capex of course, but the similar opex (maybe slightly lower as there is less physical hardware to maintain).
I seen some specialized hardware like google's TPUs. Not sure how they compare on performance per watt with GPUs though. Regardless the manufacturing processes are still the same (EUV) which is the thing that hasn't been improving. A fully optimized specialized hardware can at most deliver a single-time linear improvement (that could be very significant, for example 30% is still huge of course) and then little compared to normal GPUs.
I don't think renewable power generation is going to massively reduce costs for data centers, especially considering power transmission hasn't meaningfully reduced in cost. If anything the only thing that I think will have significant impact for data centers would be dedicated nuclear power plants physically located right next to the data center.
In fact I expect power generation to get more expensive as demand can increase faster than supply can be established. I imagine setting up new solar farms and transmission lines to be significantly harder (as in, takes longer time due to approvals and so on) than new data centers (which requires a single large location and I assume less approvals).
I don't think the main use case is for a human to directly interact with the raw token stream. You probably want reasoning and you want the thing to be able to program on its own. That uses way more tokens than you can read.
You can also look at what's been happening in mobile and especially with Apple's integrated processors. They are more power constrained, so people worried more about power there.
I think raw flops per watt come mostly from fab process, not architecture. This was my original point, fab process is not getting better at a linear (much less exponential) scale anymore.