IBM Watson's Visual Recognition Demo
visual-recognition-demo.mybluemix.net
visual-recognition-demo.mybluemix.net
For a project I've been building, I have used Clarif.ai, Google Vision API, Watson, and Imagga. From empirical tests, Watson has always provided poor, if not hilariously non-sensical, results. As an example, I use this image of a baby in a stroller (https://dl.dropboxusercontent.com/u/898689/IMG_2948.jpg). Here are the results in classification tags from the 4 services:
Clarifai: people, child, vehicle, woman, one, man, outdoors, emergency, accident, protest, adult, transportation system, carriage, portrait, wheelchair, safety, road, bike, leisure, wheel
Google: baby carriage, car seat, child, vehicle, diving equipment
Watson: performing, escalator, repairing, indoors, celebration, dancing, human, amusement arcade, bottle, baggage claim, group of people, appliance, tiger, people, big group, child, mixed color
Imagga: people portraits
These results obviously may vary depending on image subject, composition, etc., but I've basically dismissed Watson as a viable off-the-shelf visual recognition API. That being said, if you have a specific dataset of images that positively and negatively identify a concept, Watson's custom classifiers may be of interest although I haven't tried that.
From my brief interactions, they seem like awesome people too.
We provide a medical image analysis platform. I work at SemanticMD.
However, the real power of the service is the ability to train it for your specific domain - after setting up one or more custom classifiers, you can get far more accurate results on the tags that matter for your use case. You can see a few examples of this and even try it out yourself on the "Train" tab.
This isn't actual visual "recognition" but rather pattern matching, I guess.
Another way of saying this is that while the tiger training data consisted completely of tigers (for positive training examples), the animal training data might not have had any tigers at all (however unlikely) -- i.e. it could be recognizing the tiger image as containing similar enough features to the images of dogs and cats and lions that it did see to trigger the animal classifier, but with only 84% confidence.
This is the biggest gap I find in existing classifiers, the inability to make leaps from trained images (stack of pancakes) to others (single pancake) that are trivial for humans. I suspect this is because most scoring algorithms for rating classifiers don't appropriately penalize absurdly wrong answers, so what we wind up with are systems that are simply very good at matching a new image to a similar reference image, rather than really "figuring out" what the picture is of.
For example, this image of President Obama on the phone at his desk:
https://farm2.staticflickr.com/1445/24129414022_f89da8ea52_b...
--- currently is interpreted by Watson's API as "90% person"...a month ago, when I was demoing the product, it also returned "person"...but a whole variety of other things too...most notably, "Flag Burning"...it's only when you dive into the API and request a list of default classifiers did you see that the API contained a lot of specific terms, but was by no means comprehensive...so "Flag Burning" had been trained, but not just "Flag". It seems they've cut down the vocabulary at this point, which is good, because it shouldn't be judged against APIs that purport to do face/celebrity recognition and have been trained for that.
I think the key advantage of this API is that it's the only one that I had seen that allowed you to build your own classifier. Here's an example I found in the wild, of someone building a classifier to differentiate between M1A1 Abrams tanks and...non-Abrams tanks (you upload a set of positive and negative images to the API):
http://cmadison.me/2015/12/03/classifying-tanks-with-the-ibm...
Here are some quick notes I wrote for my class...I have no idea if they apply exactly today:
https://github.com/compciv/watson-preview
And here's a example of the Watson classifiers performing on sample White House images:
http://stash.compciv.org/samples/watson-preview/printout.htm...
My favorite is the pic of Obama and the Pope, of which the top classifier is "Coffee_Maker -- 0.681709"
http://stash.compciv.org/samples/watson-preview/pics/popehou...
Now...the image is classified as: 99% Person, and 80% wedding...which, actually, isn't the worst guess :)
http://builtbyswift.com/wp-content/uploads/2015/10/2015_Swif...
http://theradavist.com/wp-content/uploads/2016/03/DRIVESIDE_...
http://theradavist.com/wp-content/uploads/2016/02/morgan-tay...
It thought bird as the most probable
I've heard from teachers that children these days have an interesting new adversary with social media. On the contrary I've heard from people a bit older than me they are grateful things like Facebook didn't exist when they were teens. I'm in my mid 20s, I was perhaps the first WWW native cohort (I started using Netscape in early 1995), but social media was originally text (AIM) when I was particularly stupid, then Myspace and Facebook caught on at high-school age so I kind of straddled the two generations.
Anyway... so you could probably think of it as you might think of web development solutions, that is, there isn't one thing, but a suite of all kinds of things, and that IMHO is Watson today. I think they have a big gaping cloud offering as well to run those models they have or come up with as well, but it isn't like Watson is this single computer that won Jeopardy, and that exact computer they are now letting everyone use or something... I mean it is, but it also isn't.
This obviously requires some training