In the section "GPT4 cannot do true planning" (42:40), he asks it to change a math identity in a way that it fails to do, whereas he shows how he can do it with a little planning ahead. But this really isn't a fair comparison. If you tell GPT4 to show its reasoning before giving the final answer (i.e., to plan ahead), it also gives the correct answer.
It seems that the major difference is that we usually "don't show our work"-- our inner monologue is private.