Have you tested this approach with blind users? I think building a picture of an environment is a good task to offload to the brain and a good skill to have/develop for blind people.
> One big issue is figuring out how to sonify depth information so it's useful. One simple approach is to do a sort of sweep across each frame from left to right, letting each row of an image correspond to a certain pitch. I don't think this is a good approach, as it seems very vision-oriented and is likely to sound just like noise.
I think this is a quite good approach, but agree it has a high learning curve. However, that high learning curve might reward the end-user with a system that is more flexible. By preprocessing the input and generating audio based on the detected patterns you limit the applicability of such a system. That being said, a generic system that gives "unfiltered" output and has additional cues you can set for example for fast approaching objects might be useful.