Isn't using pretrained CNN models make embedding biased towards extracting features that were relevant from the dataset (i.e. ImageNet) it learned from? Fine-grained classification uses triplet and siamese networks to learn an embedding based on semantic similarity, but you have to define what is semantically similar in the dataset. I am curious how well pretrain networks generalize for indexing for image search. I think finding papers how pinterest and ebay apply visual search at scale may shed some light on this.