We do this with Pinecone, but we use CLIP embeddings of images, and they work incredibly well. It's kind of crazy how easy it is to get semantic search of images these days.
CLIP also does caption embeddings, so you can lookup images via both images and captions.