The semantic understanding feels much richer than diffusion based modeling, e.g. the trees on the shore growing to match the manipulated reflection, the sun changing shape as it's moved up on the horizon, the horse's leg following proper biomechanics as its position is changed. I haven't gotten such a cohesive world model when doing text-guided in-painting with stable diffusion etc. This feels like it could very conceivably be guided by an animation rig with temporally consistent results.