It doesn't. Inference is still expensive, and demand for it is high, as evidenced by Anthropic's frequent "we're out of quota" messages and Deepseek's crap-out under load last night. On the training side right now only the top flight labs can conduct serious, ambitious research, and even they don't do as much research as they'd like. Witness Meta effectively train the exact same architecture on similar data mixtures for the past couple of years. More or less the same situation is happening across the board - compute bandwidth (and therefore the ability to experiment) is scarce. What this means is inference will remain quite expensive in the foreseeable future, especially multimodal and long-context inference. Believe it or not, even Google is compute constrained. When I was there some days I couldn't even get a handful of TPUs to do my job - everything was allocated to training Gemini. Even if it gets a lot cheaper to train models, you could just train larger, more capable models and do more architectural / efficiency research, and iterate faster, with tremendous payback in the long run. NVIDIA is the only viable seller of shovels for this gold rush for everyone but Google and Anthropic. Bypassing the gatekeepers, and making capable AI models available to more people makes their product more valuable.