In computer science we have a way to score how good lossy data is. That way is to make it lossless and look at how much data the arithmetic coder needed to correct it from lossy to lossless.
This is a mathematically perfect way to judge (you can't correct lossy data any more efficiently than an arithmetic coder). All the entries here do in fact make probabilistic predictions on the data and they do all use arithmetic coding. So the suggestion misses a key point of CS involved here. I don't mean to be rude about it but the idea does need correcting.
Only if you're using a very particular and honestly circular-sounding definition of "good".
Some deviations are more important than others, even if you're looking at deviations that take the same amount of data to correct.
Think about film grain. Some codecs can characterize it, remove it when compressing, and then synthesize new visually matching grain when decompressing.
Let's say it takes a billion bytes to turn the lossy version back into the lossless version.
The version with synthetic film grain still needs a billion bytes or maybe even slightly more bytes, even if the synthetic grain is 95% as good as real grain. The cost to turn it lossless is the wrong metric.
It looks like LLM's are like compression algorithms with strengths and weaknesses in different things.
Losslessness doesnt always equate to usefulness. But yea, maybe a different competition for this.