How to build a simple image and primary-colour similarity database
blog.wearewizards.io
blog.wearewizards.io
https://en.wikipedia.org/wiki/Munsell_color_system
What is interesting about this color system is that it isn't symmetric (your eye is better at differentiating certain hues like blue at more maximal values than others) so more care is needed on the nearest neighbour lookups.
If you are using this coloring scheme for labelling colors in a photograph then you also need to adjust for colors that are especially vibrant and overtake the photo. For example this photo:
http://1.bp.blogspot.com/-OkTS4Wo34UM/Uwg568lRl2I/AAAAAAAAFi...
Is best tagged "pink" not "brown".
HSV is not good for perceptual comparison. Try HSP [1].
For those interested in some of the commercial solutions in the space, Cortexica ( http://www.cortexica.com/ ) is doing interesting work using neural networks for fashion image similarity. Sadly, they seem to be focusing on white-label solutions rather than having an API.
I chose upper-left simply because these sorts of photos would likely would be framed in a way that the background is visible in the upper margins, while possibly cropping the object at the bottom of the frame.
I want to stress that this code is massively simplified to avoid obscuring the main plot. There's so much we skipped: the code in the blog post actually computes pretty terrible results compared to an even moderately tuned pipeline :)
That's fine. I am a day away from starting an image processing app and was thinking of Python and some libraries. This blog confirms that idea and gives the right amount of "starting point" for my effort. My requirements are very different so showing a more specific solution while still interesting is of less universal value. Perhaps a part-II with more application specific methods would be good.
I would submit that most professionally photographed fashion items would endeavor to include sufficient contrast between the background and the item being photographed.
Some of the issues that anyone attempting similar will bump into are:
(1) If you're looking at shape features, a heavy dependency on product positioning. We found you need to do classification and use secondary features to get around this.
(2) Avoiding matching on background colours. For flat-background images you can sample the corners. For gradient backgrounds you can need some kind of edge detection.
(3) Shadow removal. A lot of the colours extracted from images are shadows of the actual perceived colour. You can generally pick the perceived colour by clustering the relevant pixels by hue/saturation and picking the middle of the lightness curve bit it varies by material.
(4) Data. Fashion retail products are numerous and go out of stock very quickly so to get any real value out of this specific example you need a nice efficient data pipeline and capacity to handle constant change
(5) Query performance / index size. We used an inverted index based on Lucene to query the data. To get decent relevance you need to include more visual words than you'd typically use in a text query. In our setup we found 25-30 visual words per feature was about right. This meant you needed to keep the index size under a million records, which meant we needed a distributed index. MLT queries are useful here, but last time I looked, which was admittedly a long time ago, Lucene had an annoying performance problem that I had to work around (https://issues.apache.org/jira/browse/LUCENE-1690)
(6) Value prop. Make sure you're optimising for the right use case, because as other commenters have pointed out, the relevance of this kind of system needs to be tuned to what you want to accomplish with it and the kind of data you'll be processing.
Nobody should let themselves be put off by that list though. There's a lot of value in it, and it's a lot of fun to play with. Happy hunting :)