Did somebody ever run an image of a scene vs some parts of the scene in e.g., CLIP?
Did somebody ever run an image of a scene vs some parts of the scene in e.g., CLIP?
An added benefit is that if we're smart about how we construct the lower-dimensional space, we often find that it generalizes to unseen data better than if we'd used the high-dimensional representation, because a lot of the variation we're throwing away is specific to how we constructed the input data.
A really cool thing about embeddings is that we've known how to do this for a really long time --- for example, a landmark paper [1] on this exact problem was published in the 1930s!
[1] Eckart, C. and G. Young (1936), "The approximation of one matrix by another of lower rank." Psychometrika 1, p. 211-218. https://link.springer.com/article/10.1007/BF02288367