What do you mean? Is post-training/RL not training? It seems to me that inference is quite stable there and only change being more time on RL which I would still categorize as training.
You got it backwards. RL is long running inference plus training at the end. They let the Agent spin for quite a while, then they calculate the reward afterwards and modify the parameters accordingly.
... so training? Even "regular" pretraining is "inference plus training at the end" if you want to frame it that way