Top contestant in the Hutter Prize uses a neural network for compression. So fair to say, LLMs would perform pretty well compared to gzip.
Even ignoring speed per GP, the Hutter Prize's metric includes the size of the decompressor. LLMs would be disqualified for being larger than 1GB.
Hutter prize does have speed restrictions. If it did not, LLMs would win even with counting the size of the model (which is the most reasonable choice imo) as per the main benchmark: https://www.mattmahoney.net/dc/text.html.
And the hutter prize disallows GPU's. If you allow use of a powerful GPU, you can do quite a bit better.
hallucinations are lossy compression artefacts
can they be considered to have compressed the entirety of their training data into their weights?
If you consider them lossy compression, then yes.
The goal would be to find the minimum model that, with a fixed seed, would exactly reproduce your text.
So I could hide information in a model basically
I think this would be the exact opposite. This model would be optimized to only output your information.
thats what I meant. For a given input give me the output I want without anyone knowing why that is there
lossless vs lossy is the question.
How lossy? Because I can lossy compress anything into 0 bits.