Huh. This seems to boil down to 'noise is higher information entropy than realistic content; partial learning will learn realistic content before learning noise' or something like that.
I think that can be mainly attributed to the fact that the last few deconvolutional features are overfitted to features in the image and are somewhat robust to noise. The network does not even learn features to produce e.g. white noise as output. This is probably much less magical than the paper makes it seem to be.
Interesting. Then you might achieve similar results by compressing the image as a JPEG with low quality settings.