Building a natural description of images
googleresearch.blogspot.com
googleresearch.blogspot.com
I imagine I might make similar errors if I only got little jumbled fragments to work from. Given those conditions, the cat "laying on a couch" or the dog "jumping to catch a frisbee" hardly even seem like errors to me.
This is going to get radically better when someone works out an efficient way to keep the spatial relations.
1. Object recognition (there's dog and frisbee in the photo) 2. Object localization (Dog and frisbee's ROI in the photo) 3. Relation estimation (Based on X factors, the dog might be chasing the frisbee).
Not sure what you meant by spatial relations (localization?) but recognizing (what) and localizing (where) would be key to drawing relationships between objects.
Really impressive work but definitely not a leap.
http://www.huevaluechroma.com/pics/3-4.jpg
This is also true to some extent with the actual hue (i.e. red versus green, rather than brightness; see image C) but less so.
So, it knows "pink". And the motorbike isn't borderline pink / red -- it's not like hunting pinks -- it is definitely pink.
Having said all that I'm amazed at the results. It feels like I'm living in the future.
Just to be clear, I don't mean it as a criticism, it just seems to be the easier part.
Similar error - yellow passanger car is described as yellow school bus - school bus is more common in yellow color.
That said, I don't expect particularly high accuracy from the composition of an image recognition system and natural language generation. The first actual demo of this is going to be a source of utter hilarity. I hope they're okay with that.