The test is rigged because they used non thinking models.
But also:
GPT 5.2 Thinking, Standard Effort: Walk - https://chatgpt.com/share/699d38cb-e560-8012-8986-d27428de8a...
I'm assuming "GPT 5.2 Thinking" is, in fact, a thinking model?
If you ask GPT 5.2 with high reasoning efforts in the API, you get 10 out of 10: drive.
And the problem is NOT that I'm using a product in the advertised, intended way.