For instance, consider the distribution of strings drawn from the language '0+'. Now consider the same for the language '[01]+'. A compressor looking at only the strings of the first language within those of the second can do a much better job if it does not have to account for future data.
This also relates distantly to the idea of overfitting in machine learning.
But the Hutter prize (of which I'm a big fan) is for ever-more-optimal compressors, and in fact many of the solutions don't generalise to other input data without stripping out various tricks.
Most of the neural compressors are trying to squeeze every last bit out by sacrificing runtime. You can speed them up a fair bit and they're pretty good. fast-cmix with a small priming corpus will probably be about twice as compact as something like lzma or brotli.
Eg: cigarette smoking was extremely common and a recommended social activity in the past, but now it's not.
If you asked someone would they still smoke if cigarettes were harmful, and they tell you that it's an out of distribution question because they don't know a world were cigarettes are harmful, you'd question their intelligence.
We want predictive models to be able to estimate the consequences of different (counter-factual) actions to then pick the best one.