I haven't read the paper in full detail yet, but I do have a minor editorial comment: while the appendix L.2 was satisfactory, I thought the condensed argument in 5.2 was a bit too sloppy. In particular,
H(X) + H(Y) = H(X | Y) + 2I(X ; Y) + H(Y | X)
By discarding H(Y | X) - which appears again when predicting at the following position - we observe that 2-token prediction increases the importance of I(X ; Y) by a factor of 2.
The argument about "discarding" was not clear to me - if you're predicting the third token Z, then shouldn't H(Y | X) be contained in the implicit context C, and therefore can't be freely discarded? I don't think this argument was clarified in the appendix. But this is mostly about presentation, I wasn't so confused as to doubt the gist of the argument.