CaptionBot by Microsoft
captionbot.ai
captionbot.ai
I don't mean to offend, but I'm left wondering if the creators of image recognition services disincentivize their neural nets from recognizing something as an ape, gorilla or chimpanzee so as to avoid the same mistake Google made when it falsely recognized black people as gorillas [1].
[1] http://blogs.wsj.com/digits/2015/07/01/google-mistakenly-tag...
Being mislabeled as a bear on the other hand does not generally carry the negative connotations as calling people apes, so I suspects that this is why CapitionBot is erring on this side of the classification.
I suppose it's like calling someone a Neanderthal, but with racial implications. Interestingly, we're learning Neanderthals were not as cognitively lacking as perhaps some supposed.
It's like saying, hey, you are hairy... Right, we're all hairy (bald or not) with very few exceptions. It's a weird thing.
[1] In Indonesian lang. people and apes share the "orang" root. Oran, orangutan, orangnakal, etc.
It may well be taxonomically correct to say that humans as well as apes are primates and members of the Hominidae family, but that doesn't make it any less insulting.
Just because something is technically correct doesn't mean it isn't dehumanizing and suitable as an insult. I'm sure nobody would like to be called a sack of meat either.
Still, this is a machine that's doing the labelling, and one would have to think someone purposely programmed it to do that, rather than the engineers lacking foresight to see a possible misrecognition.
Wow. 10 years ago it would have been seen as a comical blunder of a stupid AI. Something to be fixed for sure, but not a Serious Social Issue by any means. Nowadays, it apparently warrants several follow-up articles in the mainstream media, "social" commentary and 3,309 retweets.
From the blog post: “The bias of the Internet reflects the bias of society,” she said.
In some cases - yes, but this one seems more like society deliberately projecting human motivations onto a primitive algorithm and jumping to conclusions about what its errors "really" mean.
-
For people who disagree, here is a scenario your might want to consider. Imagine that you built an image tagging service. Imagine that someone found a glitch in your service that they consider offensive. Imagine them tweeting about it (before or instead of contacting you directly) and getting a similar kind of reaction, complete with extensive social commentary and media coverage. Nice, big crowd of people using your company and your service as a convenient example of things-that-are-wrong-with-our-society. How would you feel in that case?
I'm all for not discriminating people, but on the other hand I firmly believe that technical progress should not be second to people's feelings.
Witness the KFC ad (https://www.youtube.com/watch?v=ZaIhf41ctkM) which was broadly labeled as "racist", despite the fact that the 'black people love fried chicken' stereotype is (as far as I know) only a U.S. construction.
I understand, but isn't the logical conclusion that one cannot make a piece of technology (like an image caption bot) that is unaware of, for instance, such complex racial relations? And if so, should technological progress really be hampered by people's sense of outrage, even in the absence of malicious intent?
I don't think it's reasonable for anyone to expect an AI system, at our current level of technology, to have a complete understanding of every nuance of human pique to the degree where it will never do anything which could be interpreted as offensive by anyone.
Hell, that's a far higher bar than we humans can hope to meet in today's 'outrage culture'.
Is this assumption based on something specific? I.e. are there good reasons to believe that they weren't trained on a diverse set of people and that such training would prevent this error from happening?
>I am not really confident, but I think it's a man is smiling for the camera and they seem . I am 99% sure that's Pope Benedict XVI
Source Image: http://memesvault.com/wp-content/uploads/Wat-Meme-Old-Lady-0...
Needless to say, my errant habits of trying to break stuff shine through once again.
>I am not really confident, but I think it's a man holding a camera.
Source Image: http://www.cinemablend.com/images/news_img/79237/pulp_fictio...
While we're here, let's go the full way and set up a proveable and public way to train a robocop, and I'd trust that more than a human cop. The awkward moment when AIs have more brains than cops (at least under the US system).
So Jane truant with convictions of petty theft or battery gets off with probation if she agrees to embed her own personal "robocop". Yes invasion of privacy, etc. But the alternative for her would be time in the pen, for example. So in this case people can become their own robocops who turn the host in to authorities if certain conditions are met (engages in previously restricted activities).
I think this is more likely than a roving robotic cop which looks out for misdeeds.
Edit: They should have named it CationBot.
https://www.microsoft.com/en-us/legal/intellectualproperty/c...
He didn't mistake the football player with a rugby player, a cricket player or else. And +1 for the emoji
Wonder if it was the ball logo on the player's shirt, the greater popularity of football , or something else that helped it distinguish.
I uploaded a photo of the planet Saturn and it guessed that it was dish. It got the shape right.
Wondering if you plan to open up a caption API of any sort? Can definitely use something like this. If you desire the training feedback, then that could be added as well as part of the API. I'd be willing to do that for some images. So if you do add a training feedback API, please make it optional.
For example, I uploaded a picture of my daughter as an infant, and it said, "This is a baby on a bed and he's :D" which I gave four stars to because it said he instead of she. But honestly you really have no way of knowing that. :)
It seems from the comments that that is one of the most represented misclassifications.
CaptionBot seems to have a bit of trouble with simple two-colour outline drawings. In one case I saw it even get the colour wrong ("red and white" for a black and red image).
Is that something you would have expected?
Also- I notice it doesn't do very well with character recognition either. Is that surprising?
https://news.ycombinator.com/item?id=8129499
(Site's gone now but it was a demo on top of Clarifai)
Honest question, is this much better than Clarifai?
https://www.dropbox.com/s/ty34c02y1mngyrc/Screenshot%202016-...
and I got this result "I am not really confident, but I think it's a couple of glass vases with flowers on top of a surfboard."
[1] http://vignette3.wikia.nocookie.net/lovecraft/images/9/95/5b...
:)
https://upload.wikimedia.org/wikipedia/commons/thumb/9/90/Pe...
And it came back with, "I'm not really confident, but I think it's a person on a surf board in a skate park."
MS has been providing the net with fun activities for a few months now. The other night the celebrity look alike site blew up on twitter and instagram
"In a development that surprised even us, after an influx of /b/ and Something Awful Goons, the AI decided to shut itself off."
Edit: recognizes Stalin though http://i.imgur.com/9W6wqUd.png
The caption guess was "I am not really confident, but I think it's a television screen and he seems ."
http://i.imgur.com/kS6sgNT.png
Edit: I was expecting it to think an eel was a snake, but... http://i.imgur.com/EmpRNkA.png
It is far from perfect, but is near state-of-the-art. I'm guessing it won't hold up to HN.
On the other hand, it described a picture of Grand Prismatic Springs in Yellowstone as "a train with smoke coming out of the water." Which also is kind of like the crazy things kids sometimes say when they see something new.
I get the same "about as smart as a very young child" vibe from CaptionBot; the intersection of pleasantly silly and rather impressive.
(Photo: Rover P6 on the start line with a Nissan R35 GTR, guessed as "I think it's a police car parked in front of a truck.")
The last one in this set really surprised me:
In fact, I assume this is a crowd sourced training for the tech..
Kind of disappointing, but at the same time I understand that this task is not trivial at all.
As far as I know this was the first research to do the super cool thing to combine multiple neural nets trained on different data in super cool ways:
"Now, what if we replaced that first RNN and its input words with a deep Convolutional Neural Network (CNN) trained to classify objects in images? Normally, the CNN’s last layer is used in a final Softmax among known classes of objects, assigning a probability that each object might be in the image. But if we remove that final layer, we can instead feed the CNN’s rich encoding of the image into a RNN designed to produce phrases. We can then train the whole system directly on images and their captions, so it maximizes the likelihood that descriptions it produces best match the training descriptions for each image."
AND
"Our alignment model is based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding"
This demo uses the Vision API and the Emotion API from here: https://www.microsoft.com/cognitive-services/
Google: we cannot move forward with 'Progressive Webapps' if you guys don't fix these silly bugs.
Take-picture-do-something is a common feature of webapps like here Mr CaptionBot!!
I have a bunch of tabs open which included a video (~30minutes)
Using a Nexus 5, Android 6
1) Close-up of a roman coin
- I think it's a banana peel
2) Inverse black-on-white outline drawing of a wolf howling at the moon (logo of comic series Elfquest)
- I am not really confident, but I think it's a close up of two giraffes near a tree.
3) Red-on-black drawing of eight arrows with a circle in the middle (chaos symbol).
- I am not really confident, but I think it's a red and white sign.
4) Red-on-black drawing of a hammer-and-sickle (communism symbol).
- I am not really confident, but I think it's a picture of some sort.
5) Ltd Cmdr Data laughing, one hand on his chest, the other extending outside the picture.
- I am not really confident, but I think it's a man holding a wii controller and he seems :D
6) Germaine Greer biting the head off a barbie, while shaking another off its ponytail
- I am not really confident, but I think it's a woman eating a doughnut and they seem :D D:
7) Image of a tiny lilac octopus on a black background
- I am not really confident, but I think it's a close up of a doughnut.
8) Red-on-black drawing of an "A" in a circle (anarchy symbol)
- I am not really confident, but I think it's a lamppost
9) Black-and-white picture of actress Liv Ulmann
- I am not really confident, but I think it's a man with a stuffed animal.
10) Portrait of countess Elisabeth Bathory
- I am not really confident, but I think it's a woman wearing a hat and she seems :|
For the record, number (10) is spot on (though with low confidence, so may be just random).
A couple of days ago I think there was a post about Google doing a lot of development and research around creating systems that understand / categorize / comment / recognize images.
One thing I took away from reading about it is that Google has billions of images to train it with from all their different ventures.
Does Microsoft have access to anywhere near the same numbers of pictures?
Bill Gates sold it earlier this year to a Chinese Company http://www.ft.com/intl/cms/s/0/d6fbcb88-c126-11e5-846f-79b0e... (paywall),
The Chinese company has structured a deal with Getty to take over licensing outside of China.
However, having access to world class photography is great, but the images that (probably) will be the most interesting for Microsoft to recognize will be selfies, and other "crowd" created amateur photography and possibly memes.
I personally would think it to be cool to see if the bot could traverse the Getty collection and see if it could recognize the photographer of an image it had not seen before. Why yes, this is Leibovitz.
Hello kitty: http://i.livescience.com/images/i/000/024/750/i02/tarantula-...
I uploaded a picture of Michaelangelos David to the service to see what captionbot would say about it, and I got back a message "I think this may be inappropriate content so I won't show it."
It seems that after it generates the caption, this needs to be fed to some semantic pipe, so that a plane sitting on a book would not make sense, and try further.
After all, it really depends on the training data. If the picture of a train ticket was never seen by the NN, how could it answer correctly? How ever, it should try to reduce the answer to some more meaningfull info, for example instead of two giraffes near a tree, ideally would have said, it's a text and would attempt OCR.
[0] http://www.xperiax10.net/wp-content/gallery/cinema_x10/cylon...
It struggled with wildlife photos - a pack of arctic wolves was "a sheep standing in the snow", and penguins swimming was "a bird flying over a body of water" (close but no cigar).
And was told: "I am not really confident, but I think it's a toilet that is in the dark."
(I tried the spacex landing pictures too - it correctly identified "a boat in a large body of water" but ignored the ten-story rocket above said boat.)
-.-
> I am not really confident, but I think it's a cake made to look like a phone.
Nice try m$.