Uhm. You’re basically asking how the entire NN works. There is no easy explanation for that.
If the answer is just that the textual embeddings are also fed as simple inputs to the network, I already understand then.
So the model understands (kinda) who Bob Moog is, so when you include "Bob Moog" in the prompt, the model knows what you are looking for.
https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2...
is a hosted version, but you can download and host it yourself as well.