If you have 2 cameras observing the same object and you know their coordinates and orientations in 3D space there's a fairly simple algorithm to recover 3D coordinates from pairs of matching pixels between the two images. It should be covered in most computer vision textbooks, such as the Szeliski book: https://szeliski.org/Book/. Of course you also need an algorithm for matching pixels between the two images. There are a number of ways of doing that, also covered in the book. OpenCV probably has code that would help with some or all of this.
> For viewing I'm thinking some kind of 3D engine.
OK, that makes sense, I just wanted to be sure we were talking about the same thing.
EDIT: BTW I assume that you don't have access to the data from the drones themselves. If you did it would probably be much easier just to capture GPS coordinates from them.