1 - the robot moving behind the table leg (ie you have to do depth recognition of objects in the scene)
2 - the user's hand interacting with the artificial elements in the scene. Some code had to recognize a hand and figure out which element it was touching.
What strikes you as the hard parts of those videos besides the real-time requirement?