In the end, we only use the next-token head for generating. So which parts of the 2-token target H(X) + H(Y) are "auxiliary" in the sense that they help learning and which are "wasted"? H(X | Y) and I(X; Y) are useful for next-token generation while, by definition, H(Y | X) is the information quantity not related to the next token X. So we could say: "multi-token prediction trades the useful information I(X; Y) from H(Y) for the wasted computations on H(Y | X)". However, note that H(Y | X) is a next-token entropy for predicting Y from the prefix (C, X). If the attention mechanism allows to transfer computations already made for predicting Y|X to the next step, these computations may actually not have been wasted -- it was just pre-computations.