What Inception Net Doesn't See
abidlabs.github.io
abidlabs.github.io
For examples like the surgical mask, the ImageNet "mask" category is for masquerade masks and halloween masks, I suspect it doesn't contain anything similar to a surgical mask. Imagenet categories are a bit weird that way - they're a pretty arbitrary set of labels that were selected for use as a benchmark, not for practical use. It's got 100 different breeds of dogs, because it was a way to test fine-grained classification. It's got crane (the bird) and crane (the machine). It's got 'car mirror' but not 'mirror'. It's got 'yurt' but not 'house'.
It's great if the categories you care about happen to match the ones on the list (in both name and type, as seen in the 'mask' example). But otherwise you'll quickly run into the need to fine-tune the model on some of your own data to get the categories you want.
Also, that Roomba really does look like a CD player.
Of course, the big trend now is CLIP/ALIGN https://openai.com/blog/clip/ https://arxiv.org/abs/2102.05918#google which are trained on web images on such a scale that they do quite fine on anime images. Look on Twitter or Reddit for 'Big Sleep' examples using the ThisAnimeDoesNotExist.ai anime model: CLIP can generate anime images reasonably well, along with everything else it can do so well. (The 'blessings of scale', I call it: https://www.gwern.net/newsletter/2020/05#blessings-of-scale Just train on as much data as possible, and a lot of these weaknesses fix themselves.)
To try to turn this comment into something only very slightly more interesting, I find it a bit of a challenge as a career thing, not being an expert in ML but being interested, to sort through who is actually doing the worthwhile work and who is just spouting buzzwords and regurgitating truisms. But I guess that’s part of the learning as with any field.
BTW, it would be interesting to see how OpenAI's CLIP does on this problem, since it has more varied training data.
I see this as a very funny way to recall what we always knew. Like a court jester reminding the king he was always naked.
Sure, you are right, but lots of people are wandering about saying different!
Also, it may seem minor, but as someone who has been working professionally with neural networks since 2013 or so, there is a huge difference between saying "what inception net doesn't see" and "what inception net trained on image net doesn't see". I may grant that the author is simplifying for an inexperienced audience, but if that were the case, they should at least point out how such problems can be avoided, and have been avoided for at least 5 or 6 years now, if not longer.
Joke aside, humans make mistakes too, just different kinds of mistakes. Humans get tired, bored and when we die, experience dies with us as well. AI worth is based on application, for example I would train it with fruit images if it were a fruit picking robot and it would be useful, even if it can't recognise upside down cars. Like any tool you got to know how to wield it.
Humans are expert only in the general things about life and maybe a specialised field. Most humans are bad at most tasks. Try to get a quadratic equation solved by random people on the street, like x^2 - 5x + 4 = 0, without paper and pen. It's a matter of being "in distribution" for us too. Nobody's perfect at everything (no free lunch).
What's missing for our AI to be truly great is the part about evolution. AI's need bodies, environments and an evolutionary program. Supervised models can't experiment new things except for their training data, they can't modify their environment in any way, can't formulate and test hypothesis. When they will be in the world like us, they will be much better. They will be the proverbial scientist Mary of the "knowledge argument" that sees colour red for the first time, after having studied everything about it in theory.