GPT also uses embeddings. It converts each token into a vector that captures the meaning and context of the word. Related tokens are close by in this large vector space.
The way I understand it is that the best way to predict the next words in a sentence is to understand the underlying reality. Like some kind of compression.