Those models don't have any understanding of physics, they just regurgitate what they see in their vision-based training set, just like any image or video generation model does.
Monkey see other monkey cannot go through wall, monkey don't try go through wall.