Actually, of the examples they showed, all but one clip featured both camera and in-camera motion. Granted not a lot of the former, but according to my non-expert opinion, maybe enough to construct a disparity map.
I imagine having stereo video would also help generate a depth map from disparity?
Yeh, I got my terms confused. My bad. Disparity is, I believe, only possible to get from a stereo pair. TFA presents something closer to a camera track, albeit a very short one. From a camera track it is a short hop to extracting a point cloud. What make the author's approach possibly unique is that the foreground object is in motion as well as the camera. In standard VFX foreground object motion is usually avoided.