Also note that they are using entirely different methods. Microsoft's approach analyzes the content of individual frames to model the camera's position and rotation in 3D space while Instagram uses the iPhone's gyroscope.
[1] https://news.ycombinator.com/item?id=8227321
[2] http://en.m.wikipedia.org/wiki/Hyperlapse
via https://news.ycombinator.com/item?id=8227330
edit: On a side note, I'd like to point out that each approach has very interesting merits:
Microsoft's method can be applied to any video file, but requires more substantial CPU usage. Since instagram isn't analyzing video content, but instead cropping the input video based on gyroscope data, it requires much less processing time and can be done on the fly on a mobile device.
Microsoft's method can even render an entire scene from its 3D model (see their mountain climbing example where the helmet-mounted camera looks in an entirely different direction) while Instagram's method crops input video frames based on gyroscope data. This makes Microsoft's method more robust, but much less practical on a mobile device. Each method makes different trade-offs.