If you think about how a language model works, it's not totally clear that it should work at all.
Humans use sentences to communicate pre-existing ideas - ie. we already have the object, subject and action in our heads prior to forming the sentence. In order to generate realistic sentences, an LLM needs some ability to plan ahead, so that when generating the logits for the first token, it needs some idea of how the sentence will end.
But if you think about it, nothing about next-token prediction should necessarily lead to this planning ability. It can easily get into a feedback loop of predict the same token over and over (and small LMs do in fact do this) We also didn't tell it explicitly that sentences should have objects, subjects and actions, these behaviors are purely emergent from the task of next token prediction.
In a similar vein, I think artificial agents would not need explicit goals or affordances, just a sufficiently complex environment that a large model would not overfit on.