Just a fun technical note:
CLIP would be possible with any language and vision encoder.
It turns out that transformers were the most efficient encoder choices available at the time but most likely the approach would have given interesting results with a resnet and some sort of convolutional language encoder. In fact the paper has a resnet as one model for the vision side.