Except the LLM generalizes its encoding to all english text where as the copy of wikipedia can only 'compress' wikipedia.
> Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights.
then.
If all you care about is compressing Wikipedia, but ignore the size of the actual data, what is it that you are actually trying to do?