If it doesn't pattern match to that specific problem, then it's possible to have it "reason" through a similar setup:
This one I used objects A, B and C, where A and C must move together with B.
https://chatgpt.com/share/c3f744e4-199a-4c45-a332-c40d505326...
Do you think that is also from a trained dataset?
Because I wonder if it might have some sort of internal router in mind where certain signals will trigger it to pattern match to a very common problem whereas if it's not overfit there it will try to fallback to some form of "reasoning".