I don't understand how that's even possible with a "next token predictor" unless some weird emergence, or maybe i'm over complicating things?
How does it know what the next street or neighbourhood it should traverse in each step without a pathfinding algo? Maybe there's some bus routes in the data it leans on?
Also, algorithms you learn in CS are for scaling problems to arbitrary sizes, but you don't strictly need those algos to handle problems of a small size. In a sense, you could say the "next token predictor" can simulate some very crude algorithms, eg. at every token, greedily find the next location by looking at the current location, and output the neighboring location that's to the direction of the destination.
The next token predictor is a built in for loop, and if you have a bunch of stored data on where the current location is roughly, its neighboring locations, and the relative direction of the destination... then you got a crude algo that kinda works.
PS: but yeah despite the above, I still think the emergence is "magic".
Because Transformers are 'AI-complete'. Much is made of (decoder-only) transformers being next token predictors which misses the truth that large transformers can "think" before they speak: there are many layers in-between input and output. They can form a primitive high-level plan by a certain layer of a certain token such as the last input token of the prompt, e.g. go from A to B via approximate midpoint C, and then refer back to that on every following token, while expanding upon it with details (A to C via D): their working memory grows with the number of input+output tokens, and with each additional layer they can elaborate details of an earlier representation such as a 'plan'.
However the number of sequential steps of any internal computation (not 'saved' as an output token) is limited by the number of layers. This limit can be worked around by using chain-of-thought, which is why I call them AI-complete.
I write this all hypothetically, not based on mechanistic interpretability experiments.
Note also that sequential computations such as loops translate nicely to parallel ones, e.g. k layers can search the paths of length k in a graph, if each token represents one node. But since each token can only look backwards, unless you're searching a DAG you'd also have to feed in the graph multiple times so the nodes can see each other. Hmm... that might be a useful LLM prompting technique.
So if you wished you could implement a transformer by recomputing everything on every token. That would be incredibly inefficient. However, if you're continuing a conversation with an LLM you likely would recompute all the state for all tokens on each new user input, because the alternative is to store all that state in memory until the user gets back to you again a minute later. If you have too many simultaneous users you won't have enough VRAM for that. (In some cases moving it out of VRAM temporarily might be practical.)
The answer is always ‘turn left’ no matter where you are on park avenue. These kind of heuristics would allow you to build more general path finding based on ‘next direction’ prediction.
OpenAI may well have built a lot of synthetic direction data from some maps system like Google maps, which would then heavily train this ‘next direction’ prediction system. Google maps builds a list of smaller direction steps to follow to achieve the larger navigation goal.
It isn't. It failed.
> I wouldn’t use it as directions, since it sometimes got left and right turns mixed up, stuff like that, but overall amazing.