Playing for Data: Ground Truth from Computer Games
download.visinf.tu-darmstadt.de
download.visinf.tu-darmstadt.de
It also claimed they use communication to the GPU, but none of that is visible in the demo. It looks like a magic wand (like Gimp's) that selects pixels of similar color for videos, except much slower.
And finally the times mentioned: the first two images took an hour or more to label, the third seven minutes. I'm guessing that's their innovation but I'm wondering what object recognition program takes more than a few seconds to process a frame in the first place. They mention being 'pixel perfect' but any object recognition would be, given it can recognize each object in the image and thereby classify each part of the image.
Now they also create a dataset, but instead of recording and labeling the real world, they take images from GTA and use extracted mesh/texture/shader ids to automatically label all objects in an image.
However, the game does not provide any of these 'rendering resource to object class' associations by default (at least not at the level they are intercepting the game/gpu communication). So someone has to make this annotation in the first place. That is the 'magic wand' tool, where someone is still annotating, but the human effort is reduced by nearly 3 orders of magnitude (7 seconds per image) compared to the conventional way of creating those datasets.
The authors propose to just use <some open world game> to take a huge bunch of images. Since we're talking about a game, the computer has a perfect internal representation of entities and hence things that can be considered cars, trees, streets, etc. We can thus not only obtain an image per frame that looks close to the real-world, but immediately also one that is labelled.
Why is this helpful? To train computer vision models such as the ones used in self-driving cars. Of course, the assumption here is that the imagery obtained from a game is close enough to the real world, so that a trained model would continue to work in the real world. I haven't read the paper in full, but the authors experiments show that this is the case. They still use some original imagery though, so perhaps it's not possible to use game-imagery alone. I also don't think an experiment was performed to see if this method would still hold up when using games having older, worse looking engines (it would be interesting to see whether deep models could still generalize towards the real world from this).
Finally, the authors spend a lot of hacky efforts in forcing the game to outputting labelled images. As others have suggested here, they probably would have been better off contacting some mod authors (who could whip this up in a day, probably) or even the game developer itself (though I don't think Rockstar would be particularly interested to collaborate on this).