"OMG, this is amazing, it's just like humans. We're probably close to AGI."
"Ha-ha, humans are stupid, so the algorithm giving unexpected result is just a proof that it's better and less biased."
Here, we have both in response to the same demo.
Still, I honestly don't know why some people are so biased in favor of neural nets and have zero interest in edge cases and flaws (the most interesting parts if you want to gain deeper understanding of how the algorithm actually operates). Wishful thinking, I guess.
Multistable/Bistable perception is not unique to Google Cloud vision - it afflicts humans too.
I'm looking forward to ReCAPTCHA asking "Click the images of things that are upside down."
What if it was AI looking at bacteria? Or scanning a roadside for IEDs? Or when a guy on a bike when turned and rotated the correct way appears to be a crosswalk paint mark of a guy on a bike?
If our current AI is making different “DEFINITE” determinations based only on image rotation - there is a problem.
The image is both a rabbit and a duck regardless of orientation, capturing the object it depicts as a single class with a confidence measure is a mistake.
The magnitude of this mistake becomes apparent when you connect it to real-world decision making, and it becomes highly unsafe.
As is most of ML because it only uses statistical (rather than causal) modelling of the world -- so really, it is only offering us generalised statistical associations. It cannot cope with statistical discontinuities.
The trouble with the world is that the stupid are cocksure and the intelligent are full of doubt. - Bertrand Russell
Though in this case, stupid/intelligent are probably overly harsh. Intelligent people are certainly not immune to this 'trap'. In some ways they can be even more susceptible since they may themselves know very little, but that very little is still enough to put them ahead of 80% of the rest which can yield unjustified confidence. So let's just say uninformed/informed.
Many things belong in distinct groups depending on their orientation or other physical attributes.
A bucket is, amongst other things, a bucket when standing flat on the ground with the hole facing upwards.
Turn it around, and it becomes the cover for a mole trap, place it on someones head, it becomes a rain cover. Mount a light bulb inside it and it becomes a lampshade.
It's still a bucket, but it's not _primarily_ a bucket in all cases, and shouldn't necessarily be classified as a bucket in all cases. Rotate the plus symbol 45 degrees and you got the letter x. And it is definitely, 100% certainly an x in one orientation and a + in another orientation.
In most fonts, the only difference is rotation and/or mirroring.
Yet they represent different sounds, they have different meanings.
Rotation is another data point, not something entirely independent from the data.
If you showed me the rabbit rotation of the picture, i'd tell you with pretty high confidence that it's a rabbit.
If you showed me the duck rotation, i'd tell you with pretty high confidence that it's a duck.
That's the point of this, it's an illusion.
And it did give a bit of an "I don't know" answer for many of the rotations in the middle of the gif/video, which is exactly as I would expect it to, and when I pause the video at those points and glance at it, it doesn't look like much of anything to me either.
I still fail to see how human brain only fall for this illusion once.
Eg clever animal camouflage. A caterpillar that looks like a snake should not be classified as ambiguous. It should always be caterpillar.
[1] https://en.wikipedia.org/wiki/Scale-invariant_feature_transf...
Anyway, the early layers of an NN should be performing an encoding that creates scale and rotation invariance though, so that later layers can classify. That's what makes this result interesting. Well that and the ambiguity matches the human ambiguity.