I built a project at school using a Kinect in 2011. Similar UI/UX model where you can essentially 'sense' the skeleton - we routed that data through a Node app that allowed you to swipe in the air to move through a photo album.
One of the hardest parts of this is getting what people call the 'clutch' right. Basically, how is the interface supposed to know that my arm movement is meant to target the device, and not say a friendly wave to my neighbor? In voice interfaces this is the equivalent of 'hey Siri' or 'Ok google'. With skeleton sensing interfaces we could still do an audio clutch if needed or you'll need to use another body part to engage the action on the device. Fascinating problem and I'm curious as to what clever solutions will surface.