#1: it does not require deep world knowledge, because that's not what local models are for.
#2: it directly attacks drive-by understanding, overly linear processing training, poor attention mechanisms, poor reasoning patterns or lazy assumptions that ignore very easy low hanging fruit.
#3: it requires solid instruction following in the face of errors. a lot of models will run into errors and then fall back into some kind of error recovery process that bypasses instruction following.
#4: does not require prompt fine tuning to tweak to each individual model. they all seem to understand.
#5: not unfair. almost every model demonstrates in their reasoning that they have the necessary information that if reasoned about appropriately, could arrive at the correct answer.
#6: not designed to add unnecessary complication that it is intended to exhaust reasoning budgets of any sort, so it is not inherently unfair to models that reason a little more or less. for example, it does not require unnecessary reasoning soaks (ie: hiding the prompt inside base-64 encoding)
#7: has real world use and is probably applicable to overall ability to generalize.
#8: can be scaled up as models get better.
#9: is a very good indicator of how bad a model is falling apart under various inference settings.