> plus data on the camera angles they were taken from
Doesn't seem like much of a stretch to determine the angles as well.
E.g. a semi brute forced way with GANs
Doesn't seem like much of a stretch to determine the angles as well.
E.g. a semi brute forced way with GANs
It all works pretty well. Trying it on your own video is pretty straightforward.
If you're genuinely curious, look into structure from motion, visual odometry, or SLAM.
Not really, with SLAM there are various algorithms to keep inaccuracy in check. Basically it works by a feedback loop of guessing an estimate for position and then updating it using landmarks.
If you want to experiment, take a bunch (~100) of photos of an object, and use COLMAP to generate the poses. COLMAP implements a global SfM technique, so it will be very accurate but very slow.