In a nutshell, CLIP is a multimodal model that combines knowledge of English-language concepts with semantic knowledge of images. It can just as easily distinguish between an image of a "cat" and a "dog" as it can between "an illustration of Deadpool pretending to be a bunny rabbit"[6] and "an underwater scene in the style of Vincent Van Gogh"[7] and any other concept you can come up with[8] (even though it has definitely never seen those things in its training data[9]).
This is how these CLIP+VQGAN notebooks can create such a symphony of artistic renderings[10] (CLIP steers the GANs towards its interpretation of any English string that the artist can imagine).
[1] ELI5 CLIP: https://blog.roboflow.com/clip-model-eli5-beginner-guide/
[2] How to Try CLIP: https://blog.roboflow.com/how-to-use-openai-clip/
[3] Content Moderation with CLIP: https://blog.roboflow.com/zero-shot-content-moderation-opena...
[4] CLIP Prompt Engineering: https://blog.roboflow.com/openai-clip-prompt-engineering/
[5] CLIP for Semantic Image Similarity: https://blog.roboflow.com/apples-csam-neuralhash-collision/
[6] Deadpool Pretending to be a Bunny Rabbit: https://paint.wtf/ranking/erFA7/uCzEnPBtKkgoBAhBesQY
[7] Underwater Van Gogh: https://paint.wtf/ranking/xboNk/0fjxp0kEi0WeDgkP3LZX
[8] Other human drawings as judged by CLIP: https://paint.wtf/leaderboard
[9] CLIP-judged Pictionary game: https://blog.roboflow.com/how-we-built-paint-wtf-an-ai-that-...
[10] AI Generated Art with CLIP+VQGAN: https://blog.roboflow.com/ai-generated-art/