I wonder how far you could go generating tons of images from human art "feedstock", having people select the best 0.1% of them, then using them as input for another round, and so on. Then replace the people with another AI whose job it is to grade aesthetics of an image somehow, for a fully closed loop.
Edit: In other words...
1. Train your stable diffusion model using a corpus of normal human images.
2. Use the model to generate thousands of images.
3. Use human analysts to train another AI how to find "good" images in that set.
4. Generate ten million stable diffusion images.
5. Use the other AI to find the "good" ones.
6. Use those to train another stable diffusion model.
7. Goto 4.
(You lose some information each iteration, though, so it would probably eventually converge to grey goo or millions of pictures of Agent Smith.)