The "reward signal" aka objective function seems to be the challenging part here. In the parent post's suggestion, all you'd end up with is a neural network that could /maybe/ reproduce a picture (assuming CSS was capable and the network has the approximately learnable properties necessary). It'd be more interesting to have some "quality" measure that actually meant something to evaluate outputs.