I am not very interested in whether or not the dataset that these authors chose is a pleasing dataset to use as a basis for a video game aesthetic. I am interested in in the technical accomplishments that this model was able to achieve in regards to preserving image features of semantic importance (e.g. writing, symbols, logos), not "hallucinating" new image features, and maintaining frame-to-frame stability in the output video product. These are very notable accomplishments for an adversarial model in this space, and the techniques they used to accomplish them are interesting. That Rockstar could have made their game more Cityscapes-like without this tech is missing the forest for the trees.