In the Deep Vision world we as a group are trying to segment, classify and reinforce our NN training on labeled real world data. The challenge is, it's a very manual process to label data - specifically images. The more that we can do inside the computer, for example automatically labeling pixels inside an image, without having to acquire and label real world data (or making it easier to do with real world data) the easier and faster training becomes that can be applied to real world use cases.
The trick is making the virtual world match the real world as closely as possible so that the nets we make are accurate representations of real world scenarios.