I was recently trying to understand what the state of the art was here, and was surprised to learn that this is true: most 3D reconstruction fares better with a smaller number of high-resolution still photos than with a larger number of lower resolution "stills" extracted from video. I'm still somewhat confused why this is the case.
My limited understanding is that main advantage of working with stills is that they more commonly contain location tagging, while frames from video do not. Compared to the base level of difficulty of building a accurate 3D model from lossy 2D sources, figuring out the trajectory of camera doesn't seem too hard. Once one has this, wouldn't the video be just as easier to work with?
Trying to figure this out this discrepancy, I got the impression that many researchers may actually be trying to solve the harder problem of trying to compute the camera trajectory in real time, as might be needed for a self-driving vehicle: http://webdiis.unizar.es/~raulmur/orbslam/. This indeed does seem harder, but still leaves me wondering why a multipass approach wouldn't be feasible.
What am I missing? Why can't one do a first pass to calculate camera trajectory, possibly a second pass to combine temporally adjacent frames for greater resolution, then create a better model from the resulting wealth of data? Alternatively stated, why does the quality of the 3D model seem to depend more on the resolution and quality of the input 2D images rather than on the number of these images?