I was wondering how were they getting depth from a video where camera is still.
> we utilize a comprehensive set of data-driven priors, including monocular depth maps
> Our method relies on off-the-shelf methods, e.g., mono-depth estimation, which can be incorrect.