With automatic1111 you can get around this by upscaling then inpainting the spots you want more detail and specifying a specific prompt for that particular area.
The model can't just work on arbitrary image sizes because the model was trained with a fixed number of input and output neurons. For example, 512x512 is Stable Diffusion's "native size." However, there are tricks to work around this.
Diffusion models work by predicting image noise, which is then subtracted from the image iteratively until you get a result that matches the prompt. Stable Diffusion specifically has the following architectural features:
- A Variational Autoencoder (VAE) layer that encodes the 512x512 input into a 128x128 latent space[0]
- Three cross-attention blocks that take the encoded text prompt and input latent-space image, and output a downscaled image to the next layer
- A simpler downscaling block that just has a linear and convolutional layer
- Skip connections between the last four downscaling blocks and corresponding upscaling blocks that do the opposite, in the opposite order (e.g. simple upscale, then three cross-attention blocks).
- The aforementioned opposite blocks (upscale + cross-attn upscale)
- VAE decoder that goes from latent space back to a 512x512 output
At the end of this process you get what the combined model thinks is noise in the image according to the prompt you gave it. You then subtract the noise and repeat for a certain number of iterations until done. So obviously, if you wanted a smaller image, you could crop the input and output at each iteration so that the model can only draw in the 'center'.
Larger images are a bit trickier, you have to feed the image through in halves and then merge the noise predictions together before subtracting. This of course has limitations: since the model is looking at only half the image, there's nothing to steer the overall process, so it will draw things that look locally coherent but make no sense globally[1].
I suspect - as in, I'm totally guessing here - that we might be able to fix that by also running the diffusion process on a downscaled version of the image and then scaling the noise prediction back up to average with the other outputs. As far as I'm aware no SD frontends do this. But if that worked you could build up a resolution pyramid of models at different sizes taking fragments of the image and working together to denoise the image. If you were training from scratch you could even add scale and position information to the condition vector so the model can learn what image features should exist at what sizes.
[0] Think of this like if every pixel of the latent-space image was, instead of RGB, four different channels worth of information about the distribution of pixels in the color-space image. This compresses the image so that the U-Net part of the model can be architecturally simpler - in fact, lots of machine learning research is finding new ways to compress data into a smaller amount of input neurons.
[1] Moreso than diffusion models normally do
From the Paella paper[2]: "Our proposal builds on the two-stage paradigm introduced by Esser et al. and consists of a Vector-quantized Generative Adversarial Network (VQGAN) for projecting the high dimensional images into a lower-dimensional latent space... [w]e use a pretrained VQGAN with an f=4 compression and a base resolution of 256×256×3, mapping the image to a latent resolution of 64×64indices." After training, in describing their token predictor architecture: "Our architecture consists of a U-Net-style encoder-decoder structure based on residual blocks,employing convolutional[sic] and attention in both, the encoder and decoder pathways."
U-Net, of course, is a convolutional neural network architecture. [3]. The "down" and "up" encoder/decoder blocks in the Paella code are batch-normed CNN layers. [4]
[1] https://arxiv.org/pdf/2012.09841.pdf [2] https://arxiv.org/pdf/2211.07292.pdf [3] https://arxiv.org/abs/1505.04597 [4] https://github.com/dome272/Paella/blob/main/src/modules.py#L...