I'm waiting for the day this is just a library you pass an image to and it returns an array. No, not a SaaS. Then on my own pump.io, diaspora, redMatrix etc. it just works. My data, my images, my network. I'm not against the tech at all though. Neat
The classes that it can detect are from ImageNet, so that might be limiting.
Maybe someone should create a website that lets volunteers label images for this purpose.
* http://i.imgur.com/Wex6pSR.png
Which I say is pretty much what you'd need to train a net to detect people smiling (amongst other things). Of course, there are some refinements you can make to improve accuracy and presentation of the results. My point was: you should begin with datasets that are readily available, and then improve on need (and if resources are available to justify the investment).