>I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric
This is the premise of https://huggingface.co/spaces/Jellyfish042/UncheatableEval
This is the premise of https://huggingface.co/spaces/Jellyfish042/UncheatableEval
Which implies this would probably also hold for the larger models, which are sadly not included in the leaderboard.