Text Compression as a Test for Artificial Intelligence (1999) [pdf]
aaai.org
aaai.org
https://pub.towardsai.net/stable-diffusion-based-image-compr...
Funny thing is, on a per pixel basis, Stable Diffusion and jpeg are just as good. Stable diffusion looks much better, though. The reason is that when Stable diffusion lacks the information it just invents what it expects to see there. If you zoom in on the background entire appartement buildings are invented, moved or disappeared.
So, one wonders what GPT-3 could do for text compression. On a per-character basis, I would not expect any miracles. On a generally, the same, kind-of, basis, I'd would expect something special.
In my armchair speculation for the case of severe lossy text compression, the delta/diff could easily get there, making a strong compression algo based on something like GPT-3 not really practical for doing lossless.
The state of the art in neural language models was evaluated ˜5 years ago, and it was found that standard LSTM's do very well on text compression, when properly architectured and parametrized.
The main reason for excluding lossy text compression in these tests, is that there is no clear path around requiring a panel of human judges, and that evaluation now is subjective (instead of objective).
Perhaps a different route would be to task an AI to compress Wikipedia into a (graph) knowledge base, and then test these AIs on correctly answering "multiple choice"-questions. But then intelligence becomes a proxy for measuring compression, instead of here, where compression is chosen as a proxy for measuring intelligence.
Right now the biggest thing holding Hutter Prize back is the shamefully low compute and memory limits: "Restrictions: Must run in ≲50 hours using a single CPU core and <10GB RAM and <100GB HDD on our test machine." http://prize.hutter1.net/
To even begin approaching intelligence, the compute and memory limits probably need to be 100x or 1000x larger.
In a completely unrelated field, I'm seeing pretty great results with knowledge distillation to generate AI language models that fit into 20 mio float16 parameters (40MB) and can process 1GB of raw text in roughly 12 hours.
Or said in a different way: What stops me from supplying a hard-coded pre-compressed file to circumvent their RAM and HDD limits?
For most cases the size of the "compressor" is irrelevantly tiny compared to the data, but if your "compressor" is a hard-coded pre-compressed file which it copies to the "decompressor" then the few percent of gains don't outweigh the fact that you've just doubled the size that's scored. It could be useful to include hardcoded priors iff they are very slow to compute but can be expressed in a small amount of storage.