I hope he links up two or three of these to form fuller 3d models of a room.
Software-only, or would that result in too jerky of a motion capture?
Edit: Although, I suppose the "refresh rate" of the entire system would then be 20ms, which might be noticeable. Chop that down to 5ms per Kinect, and it might be viable.
Unless the 3d environment is rapidly changing, you don't need that many frames per second from each device to capture a good image.
25 fps is one frame every 40ms.
Of course, the frame-rate would suffer.