* How would human beings perform on the text-only version of the tasks given to GPT-4?
* How would human beings perform if each grid is shown on its own, making it impossible to perform side-by-side visual comparisons?
* How would human beings perform if each grid is shown on its own only once, making it impossible to perform any back-and-forth comparisons?
* How could we give LLMs the ability to "pay attention" to different parts of images, as needed, so they can make back-and-forth comparisons between parts of different images to solve these kinds of visual reasoning tasks?