There has been some prior work on using depth cameras for navigation for the visually impaired. For example, a smart cane can detect objects beyond its reach and give haptic feedback. Microsoft Research did some work with putting the Kinect on a helmet and giving audio cues for navigation (http://research.microsoft.com/pubs/184208/VisionForTheBlind....). What I'm interested in is taking that sensory input and making it less immediate by giving it a memory -- letting it build up a picture of an environment rather than needing to point a device at something in order to know something about it.
One big issue is figuring out how to sonify depth information so it's useful. One simple approach is to do a sort of sweep across each frame from left to right, letting each row of an image correspond to a certain pitch. I don't think this is a good approach, as it seems very vision-oriented and is likely to sound just like noise. Maybe if someone was using it from birth, but for relatively fast training I doubt that approach. Other approaches do more interpretation -- Microsoft's work detected faces, walls, and floors, giving each a distinct sound for greater recognition.