You could say the same thing about human programmers, but I've never heard anyone say they think that programmers "execute" Jira tickets.
> they have internal logic that assigns to the sequence of tokens in context a next token.
I don't think this means what you think it means, because it has almost no information content relevant to what we're discussing. The probabilities that are most relevant at the level we're discussing are satisfying a reward function from post-training, which approximates to:
"What is the likelihood the solution the agent is pursuing will be marked correct by the automated grader based on the full prompt and other context provided?"
It still has to predict the next token but that isn't based on a likelihood of that token appearing in a corpus of internet text consumed in pretraining. That was eons ago. Every predicted token is shaped by the probabilities of the predicted solution, which must already be very specific and shaped completely by the request and associated context that is built during investigation of the same.