nice! I'd be interested in the method they use to put everything together. My best bet is some basic structure from motion weighted by the depth sensor...or maybe it's simpler than that...
In the video it looks like you are using a volumetric representation, perhaps an Octree+isosurface extraction?
On the roadmap, but won't be in the first version.