That is a very commodify market and as such I expect power cost per query ends up setting the pricing.
Would it ever? Demand for inference (and I suppose the same applies for training) _increases_ when models get cheaper.
I suppose real risk is that one player aggressively cuts margins (say, AWS) and others follow suite. However the release of increasingly more capable, exclusive closed-weight models prevents this to an extent.
RAM production is easier to transfer same with flash
But sell it to consumers, where there is pent up demand
The only chance of memory demand going down would be breakthroughs in model size reduction.
So on the one hand you have all the sota model makers doing investor expectation management in saying wet need a slowdown for safety (may or may not be true, but also means they don’t spend on the next training cycle before IPO? Could be making their books look better too?) and on the other we have a sense that the models aren’t yet at the stopping place where we can just not train another cycle and use distillations of the current generation?
Source/citation?
The amount of money spend on the AI hype train is crazy. Meanwhile we're wasting fab time making chips that might never be powered on. It's absolutely insane that there are more money to be made on hardware for AI that may never be used, rather than producing a product that consumers and businesses need right now.