It isn't machine learning, it's aerial photography shot at an angle with depth sensing.
For example, take an airplane and mount two cameras on it, one on each side, angled down at perhaps 45 degrees to capture ground imagery in each direction.
Add a LIDAR device next to each camera to capture a "point cloud" of the same area the camera is imaging.
Now you have photographic images with distance data for each pixel. You can use that to construct a 3D image that can be viewed from various angles.
(Disclosure: I work at an unrelated Alphabet company but have no personal knowledge of any of this, it's just my semi-educated guess.)