The Hutter dataset is included, incidentally, and the top is a Transformer-XL (277M parameters) at 0.94 BPC. Does all of that seem 'very much not AGI'?
A decompressor with a high upfront cost (say a pretrained model) and otherwise higher compression ratio would be more interesting than tweaking yet another PAQ8 variant, but I'm not sure this contest's rules make that very viable.
https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
If you think of intelligence as distilling the essence of the universe around you and building on top of that, at least the former starts to sound like really really really sophisticated compression.
I run an non-artificial GI in my brain and can't solve this challenge manually. So AGI doesn't seem sufficient.
If you studied the text and learned everything in it, you'd be able to generate a new text from scratch, from your internal representation of the information. It wouldn't be word-for-word identical, but that's actually evidence that there is a compressed representation in your head. If you got a good sense of Wikipedia's style out of it, your reconstruction would be even better. Now imagine that we let you take notes - little reminders of unusual turns of phrase, hard to remember dates, small outlines etc. You could probably do quite well indeed, perhaps even well enough that your output could be diffed with the original and result in a comparatively small file. I think it would be fair to count that as a win for you, since the computer program is deterministic and can store the diff along with its output. True, we can't measure the size of the data in your head, but compression is definitely occurring because you would certainly fail miserably if you tried to reconstruct 100 megabytes of random binary data in a similar way.
Still, to achieve that last little bit of the way to truly lossless reconstruction without the diff-storing trick will present you with a challenge, I agree. Lossless reproduction inherently requires encoding a lot of noise, to which no pattern can be matched. It's still a valid test because, although there is a hard floor on how much the data can be compressed, a better encoder will get closer and closer to that floor. And there's room for interesting optimizations all the way to the bottom - for example, with the help of a language parser and a thesaurus one could probably encode the difference between your lossy output and the original text rather better than 'diff'...