How do training sets for something like this get built up? Creating a synthetic image of a foreground object pasted onto an existing background wouldn't seem to work because it would not capture the the complex interaction eg. The image of an astronaut walking in a field of grass.
But hand labelling would be incredibly costly, far more than image classication.
There must be an interesting trick in generating training data.