I don't think I agree this is a significant shift that is guaranteed to happen. It might happen that we will go over some sort of hump where there's less training happening than there was at the top of the hump, but who knows when that hump will be? It's such a new field and there's so many low hanging fruit improvements to be made. We could train new models for years and have steady significant improvements every time, even if there's no fundamental breakthrough developments on the horizon.
And even if there was a cooldown on new training, training is so many orders of magnitude more expensive than inference that the inference demand would have to be extreme in the face of a very unrealistically rate of training for inference to be dominant.
If you believe we are moving towards the more 'star trek' like future of AI where AI observes and interprets the world as humans see it and experience it, a massive amount of compute is still needed for the foreseeable future.
If you believe we are capping out on AI capability soon for some time, then you'll see AI as more of part of the "IBM toolkit" offered as an additional compute service and it will more likely 'fit' in our existing computer architectures.
When a single user request comes in, you just want the prediction of that single input, so no backprogation and no batching. Which is more CPU friendly.