If not the kind that aerial surveys create, then what orthography are you even working on?
And what's keeping you from fusing the video frames with IMU (and, if available, GNSS) data?
And what's keeping you from fusing the video frames with IMU (and, if available, GNSS) data?
Most phones have an accelerometer and other sensors so I'm exploring if those can be used to determine the phone's movement between frames accurately enough to help me stitch it back together. When relatively close to the subject the perspective changes so quickly that matching detected features with something like RANSAC really struggles.
I'd voraciously consume any good links you have; I'm happily over my head on this and learning/iterating at every turn. I think I've accidentally given myself a relatively hard problem because of the constraints.