I haven't used the feature, but the way it's implemented feels overly complicated, especially for something like keyword search (and not similar-image-search).
If they only use the Top10 categories in their feature vector for the documents, why don't they store these categories as tags on each documented and use standard inverted-index searching and scoring. I know the vector will express how much "beach" a certain image is, but your user-supplied query doesn't have a notion of how "much" beach the user expects, so the output can be a simple list ranked using standard term search mechanisms. What am I missing?