With the old model (and I suspect this one too) it's trained to generate from a single 'seed' pixel in the center of the image. If you erase the center of the image, that's when it completely collapses.
The model generally learns to generate each pixel from its surroundings, even if the surroundings are partially missing.