How Stable Diffusion works [1] as a whole is not really hard to comprehend at a high level - but you'll need some prereqs - probability theory underlying this is explained in Variational Autoencoders [2], then Diffusion Models [3] sort of made a really cool "deep variational" autoencoder that uses small noise-denoise steps, but largely the same math (variational inference), but they were unwieldy because operated in pixel space, after that Latent Diffusion Models [4] democratized the thing by vastly reducing the amount of computation needed - operating in latent space (btw that's why the images in this HN post look so cool - the denoising is not in the pixel space!).
[0] https://jalammar.github.io/illustrated-transformer/
[1] https://huggingface.co/blog/stable_diffusion
[2] https://arxiv.org/abs/1906.02691
If the answer is just that the textual embeddings are also fed as simple inputs to the network, I already understand then.
So the model understands (kinda) who Bob Moog is, so when you include "Bob Moog" in the prompt, the model knows what you are looking for.
https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2...
is a hosted version, but you can download and host it yourself as well.