I've described what kind of video this camera is intended to capture: https://blog.mikeswanson.com/apples-mysterious-fisheye-proje...
Any projection is bound to separate areas which could be compressed more efficiently together.
A native stereoscopic spherical video encoder could improve compression even more, since side by side views are quite similar in general.
Now that's an interesting problem to solve ! (and a very hard one probably)
Existing video formats already support this for interlacing, although you could also let inter-prediction refer to earlier parts of the same frame and get most of the benefit.
Edit: I'll certainly read the rest of the articles!
Thanks!