Maybe the world would benefit from some more well documented, open licensed training/validation data?
Maybe the world would benefit from some more well documented, open licensed training/validation data?
I still am surprised to see all the same images and diagrams in the -explanation- part of so many papers. It feels like early word processor days and everyone using the same clip art.
And what's keeping you from fusing the video frames with IMU (and, if available, GNSS) data?
Most phones have an accelerometer and other sensors so I'm exploring if those can be used to determine the phone's movement between frames accurately enough to help me stitch it back together. When relatively close to the subject the perspective changes so quickly that matching detected features with something like RANSAC really struggles.
I'd voraciously consume any good links you have; I'm happily over my head on this and learning/iterating at every turn. I think I've accidentally given myself a relatively hard problem because of the constraints.