Isn't that the same as compressing the whole book, in a special differential format that compares how the text looks from any given point before and after?
But as stated above, next token prediction is a misleading frame for the training process. While the sampling is indeed happening 1 token at a time, due to the training process, much more is going on in the latent space where the model has its internal stream of information.