Deep Learning Image Classifier
deeplearning.cs.toronto.edu
deeplearning.cs.toronto.edu
... a mountain-bike / all-terrain-bike (http://cdn2.spiegel.de/images/image-730849-galleryV9-vuuv.jp...)
... or a rugby ball (http://cdn2.spiegel.de/images/image-730849-breitwandaufmache...)
... or a bullet proof vest (http://cdn2.spiegel.de/images/image-730849-thumb-vuuv.jpg)
I guess the implementation leaves room for improvement :)
AlchemyAPI: http://www.alchemyapi.com/products/demo/alchemyvision/
UToronto
Rekognition: http://rekognition.com/demo/concept
Clarifai: http://www.clarifai.com/
I'm doing it myself, but I have a conflict of interest
Rekognition API has a similar API for all developers free.
It's reliable and very fast.
Checkout their demo page.
Rekognition:
7.55% fruit; 0.92% dinner; 0.88% produce; 0.87% alcohol; 0.84% sliced
Toronto:
50% American lobster, Northern lobster; 12% plate; 7% crayfish, crawfish, crawdad; 7% Dungeness crab, Cancer magister; 4% king crab, Alaska crab; 4% butcher shop, meat market; 4% grocery store, grocery; 4% pomegranate
I find this interesting because I thought Hinton's group had state of the art tech. Who are these people and how do they do it?
http://kephra.de/Dampf/IMG_20140620_133839_800x600.jpg <- an ecigarette, and the classifier thought its a fountain pen. Well thats not bad, I got this joke/question from humans also.
http://kephra.de/pix/Snoopy/thump/IMG_20130822_135928_640x48... <- here it thought its a speed boat ... well my boat is fast, but not a speedboat, but an sailing boat. It offered several more boat types, but not just a plain sailing boat. Interesting here is that the last suggestion of only 1% could be considered right as "dock, dockage, docking facility"
Tried some other images from the lifestyle section of my homepage, but it looks as if the system newer saw a sewing machine before as it gives "Low recognition confidence", and no tags.
The reason I find it odd is that I would expect the first example on a demo to be carefully chosen to show off the system in the best light. It would be one that has perfect or near-perfect tagging. Maybe later on, I would show the shortcomings with a tricky image like this.
That being said: even deep learning requires some sort of feature engineering at times (even if its pretty good with either hessian free training or pretraining).
The main thing with images is ensuring scaling them.
The trick with deep belief networks in particular is to make sure the RBMs have the right visible and hidden units (Hinton recommends Gaussian Visible, Rectified Linear Hidden).
Happy to answer other questions as well!
This is the implementation: http://torontodeeplearning.github.io/convnet/
Also can it tell you where in the image the identified object is?