Author here. It uses color info as well as depth for tracking. Otherwise, it'd fail if you pointed the camera at featureless geometry, e.g. walls, floors.
In the video it looks like you are using a volumetric representation, perhaps an Octree+isosurface extraction?
On the roadmap, but won't be in the first version.