Stable Diffusion's U-Net is trained to remove noise from images in latent space, which the variational autoencoder (VAE) converts to and from pixel space. CLIP embeddings are used to improve the denoising step of the U-Net by using the correlations between human language descriptions of the pixel image to reduce latent noise. Neither the U-Net nor the VAR are trained to interpolate or reproduce images from the training set; if that happened the model would be overfitted and loss would be terrible on the validation set. The VAE is trained to produce a latent space that can accurately encode and decode any pixel image, and the U-Net is trained to remove gaussian noise from the latent space.
Stable Diffusion v2 16-bit is ~3GB of data. It was trained on hundreds of millions of images (minimum of 170M in the 512x512 step alone). That leaves a maximum of ~20 bytes per image that could conceivably be a copy, which is certainly not enough to directly reproduce either the style or contents of any individual image.
There is no artwork included in Stable Diffusion. There is a semantic representation of how images are composed of varied subjects represented in the latent space and what pixel probabilities over those subjects relate to human language phrases during decoding, and finally a method to remove noise from the semantic representation, e.g. starting with a blank or random canvas and interpreting what may be there, iteratively guided by CLIP embeddings. If you give Stable Diffusion an empty CLIP embedding you get a random human-interpretable image obeying the distribution of the learned latent space.