>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.
So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...
>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]
Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.
If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.