In microsoft's agi paper, there is a test for planning they put out for GPT-4. They give detailed instructions on the constraint of a poem and expect it to spit out the constrained poem in one shot. Of course it fails and the conclusion is that it can't plan. But if you think about it, this is a ridiculous assertion. No human on earth would have passed that test with working memory alone. Even by our standards, it's a weird way to present the information but they did so anyway. This is actually a major issue for most plan benchmark assessment for LLMs i've come across.
These are the kind of problems i'm talking about. If i changed the request to encourage for a draft/revise generation process and it consistently passed these tests then it's very fair to say it can plan. That's not "hacky". It's just truth. and saying otherwise is denying real world usage for a misplaced sense of propriety to benchmarks. a benchmark is only as useful as the capabilities it can assess.
a "this is how this is presented, this is how the model performed" would have been very prudent in my opinion. Maybe these guys already covered the kind of things i'm talking about(i can accept that!)....but i can't know that if they don't tell us.