Google Cloud Vision API enters Beta
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
The new wave of vision services are amazing. There are a lot of players in this field, including IBM Watson, which has a suite of vision APIs available with similar features.
One key differentiator of the Watson offering is that we have a trainable API called Visual Recognition [2]. The pre-trained APIs are excellent and have broad uses, but it's amazing to see the results from even basic training to identify image tags directly relevant to your use case. There is a demo [3] that allows you to try it out by creating a new classifier right in the web page.
You can find some demos at:
http://vision.alchemy.ai/#demo - example images that demonstrate facial detection and identification, label extraction, object identification, and so on.
Another demo at http://visual-insights-demo.mybluemix.net/ uses the Visual Insights [1] API to identify a set of relevant tags.
[1]: https://www.ibm.com/smarterplanet/us/en/ibmwatson/developerc...
[2]: https://www.ibm.com/smarterplanet/us/en/ibmwatson/developerc...
All of the freelance sites are... well, crawling with jobs involving web scrapers. What's the business here that I'm missing?
This is what slickeals and camelcamelcamel do, though only one of those uses scrapers.
Do you have any sample code that works ?
I'm not trying to pick on Google for shutting things down; I would feel similarly if this API were from Microsoft or Facebook. It's not the first time there's been an API that I think is really cool, but was very apprehensive about actually using for anything serious.
http://www.cloudtp.com/2016/01/21/dont-let-vendor-lock-in-fe...
TL;DR: You don't get the girl if you don't ask for a dance :)
Though I admit I also hesitate to replace that API with the google offering for the one app where I actually use it. The results would probably be better, but I just got burned again from stuff shutting down, and I remember Google Reader.
But the Vision Api looks cool.
If you need to understand emotional reaction on video sources, our API can fill in the gaps not currently filled by Google's Cloud Vision API: https://www.kairos.com/emotion-analysis-api
Disclosure: I'm CTO of Kairos.com
https://www.ibm.com/smarterplanet/us/en/ibmwatson/developerc...
I think Google understands how big this could be (well, why else do they do anything I suppose?). I guess you'd just have to put faith in them. I imagine those that do will build some amazing projects from this.
So for now, my best idea is to use it to build something fun with my kids.
If only Google would adopt some sort of policy where they would let you download the dataset+code when they shut something off...
curl -H "Authorization: Bearer {access_token}" --data-urlencode "url=YOUR_IMAGE_URL.jpg" "https://api.clarifai.com/v1/tag/"
...you get the top 20 tags for that image and their confidence levels, too. Their homepage at Clarifai.com has a pretty good demo that lets you see it in a more visual way.> I'm not trying to pick on Google for shutting things down
People usually say what they are thinking. In this case, I can certainly appreciate and respect questioning what Google's actions will be in the future, but the problem at hand is that none of us can tell the future. By attempting to do so, and creating a situation where we literally believe two conflicting things at once, we get mired down in illogical arguments that end up making zero sense. Worse, when we post those illogical arguments, others get dragged into the dissonance and end up making similar arguments that make no sense and the result of that is others get lost in the mess we made. For example:
> I think Google understands how big this could be (well, why else do they do anything I suppose?)
Do we really know Google understands how big this could be, or is that just us wishing they wouldn't "shut it down later"? Both thoughts are speculative, at best.
As for me, I have no way to know if I can trust Google will leave these APIs up for as long as I need them or will not change the methods in the APIs at some point breaking my code I've written to talk to it. The latter happens to me all the time!
I realize this comment may stir some emotional responses. All I would ask is that we consider alternative ways of thinking about these illogically binding "feelings" and expand our awareness to the fact that what we really want is our software, regardless of who wrote it, to be open, transparent, trustworthy and capable of running wherever and whenever we want it, regardless of when that is.
Obviously we have a long ways to go before that statement can be a reality. One can hope though!
There is also http://www.deepdetect.com
TensorFlow Serving: https://github.com/tensorflow/serving
ReCeption (actually they call in Inception v3. Not sure where I got the ReCeption name - though I'm sure I read it somewhere?): https://www.tensorflow.org/versions/r0.7/tutorials/image_rec...
Using a SVN on neural network extracted features: http://blog.christianperone.com/2015/08/convolutional-neural...
If you want a quick and dirty version here's some Python to create a web service that calls a Caffe based Image recognizer: https://gist.github.com/nlothian/c3519adb81b3452c1938
Google's face/landmark/label/text/logo detection models are open source? Or there exist open source pretrained models?
The quality and size of the training set is (at least) as important as the machine learning tools. I imagine Google has access to a pretty big data set, along with the computing resources to process it.
Google's Inception v3 pre-trained image recognition model is open source: https://www.tensorflow.org/versions/r0.7/tutorials/image_rec...
That's the hard part because as you note this is computational intensive (the training data is actually open source as the ImageNet dataset)
There is existing code for the others part that perform pretty adequately (with the possible exception of landmark detection).
Eg:
Face detection: http://docs.opencv.org/master/d7/d8b/tutorial_py_face_detect...
Logo Detection: http://www.pyimagesearch.com/2015/01/26/multi-scale-template...
Now, making it production-quality, efficient, scalable, and the rest -- well, y'know. That's why people use cloud-based services in the first place.
But I think there's less fundamental lock-in than you think. Cloudinary, for example, will let you upload an image and get a tag out. ABBYY and OmniPage/Nuance and others offer cloud-based OCR.
I'm biased - I'm at Google this year - so take this with a grain of salt, but while I have the feeling that Google can do it better and more affordably than a small team could do it on their own, I don't think that Google pulling the API would leave people up a creek without a paddle.
So you're passing up what could be substantial benefits, Google isn't going to close it anytime soon, if they do you will have months of warning, and can easily replace it with a competitor or your own.
I played with IBM Watson visual recognition API and it didn't look like it did what I needed it to (recognize a hand drawn image of a cat for example -- it just kept labeling it only as a 'cartoon').
Bummer. At least the first 1000 images are free so I can prototype it out of curiosity.
EDIT: you could send me example images and what you need from them. I could check how much I would need to extend it to handle your case.
Is it really that expensive, though? I mean I can see it being expensive if you use it inside a product you offer for free, but if it's a commercial product, I imagine the pricing isn't that much especially if you consider that using your own infrastructure for this type of machine learning and image scanning would be much much more expensive.
But I guess that's my point, the pricing puts it out of reach of free or cheaper subscription products. Even the robot example they give in the video to achieve that functionality it would have to be taking a picture every couple of seconds or so and uploading them to the cloud - at that rate it would cost almost $5 for 15mins of use!
What you need is a custom model that is trained specifically on hand drawn pictures. You can use something like TensorFlow to build this, but getting the raw data might be challenging.
Disclosure: I work for Google Cloud
1) by using the service you grant Google use of the uploaded images. (e.g. they can use your image to increase their corpus, improve the service or use it for advertising, or use it to extract street numbers for their maps, or its always private and never stored)
2) What the resulting copyright is of the returned data. If you were to build a database based on the results, what license or copyright status this would be. Would all rights belong to me, or would Google claim rights over the results.
> 5.1 Intellectual Property Rights. Except as expressly set forth in this Agreement, this Agreement does not grant either party any rights, implied or otherwise, to the other’s content or any of the other’s intellectual property. As between the parties, Customer owns all Intellectual Property Rights in Customer Data and the Application or Project (if applicable), and Google owns all Intellectual Property Rights in the Services and Software.
> 5.2 Use of Customer Data. Google will not access or use Customer Data, except as necessary to provide the Services to Customer.
Regarding translations: https://meta.wikimedia.org/wiki/Wikilegal/Copyright_for_Goog...
It's potentially a game changer, plenty of industries have piles of scanned documents. Cheap OCR means this data suddenly becomes accessible even if the value per individual document is low (i.e. for input into machine learning).
http://www.educatingsilicon.com/wp-content/uploads/2013/10/p...
A lot better for text in photographs. Comparison might be different on dense document text though.
In training an AI system with hundreds/thousands of bits of data, no single piece of training data makes much of a difference. If one of my images on the web that I had captioned with the keyword 'dog' was used to train this system about what a dog looks like, is the model they end up with a derivative work of my captioned image? Yes, but my data would make up an infinitesimally small part of that model. Yet, in aggregate, the trained model might almost wholly rely on lots of copyrighted, rights-reserved images.
Would the resulting model be a copyright infringement? It would seem as though no rights owner would have a substantial enough claim. Yet, without all of the copyrighted works, perhaps the model would be ineffective.
Well then your ex-lover clearly deserves co-songwriting credits. As well as their parents. And anybody who has influenced them personally, or even anyone who made the food that they've eaten. There's gotta be a point at which the original is just too far removed from the end result for it to be infringing, otherwise you could just keep going.
Also is a model that classifies things isn't the same as those things themselves, or images of those things, and would most I certainly hope it would be considered transformative enough to not be infringing (I am not a lawyer though). I could give it an image of a dog, and it will tell me what it thinks it is. But there doesn't seem to be any way for me to say "show me a dog" and get back any sort of image, infringing or not.
Some thought experiments: Let's say you have a copyrighted photo, and I design an API that allows anyone to upload a photo and get a true/false of whether or not it's the same file as your photo. Is this copyright infringement if I never release the original photo?
Is deep dream original in your opinion?
Besides, it's a call without any feedback, so it's not that valuable as far as training goes.
To me this is taking copyright protection way way too far and the kind of system that could severely gum up progress and innovation.
- Would you primarily use the plugin to auto-tag uploads in the admin area? - Would you have problems with getting your own Google API key (and thus billing via Google), or would you expect the plugin to take care of this as well? - Any other expectations?
As said, what I have now is not really release-worthy, but I can get it to that point in a matter of days.
Thanks.
- Aron
"TEXT_DETECTION 1024 x 768 OCR requires more resolution to detect characters"
How big were the images you used?
While this technology is fascinating, I can't help but feel a little unsettled reading that.
http://developer.affectiva.com
disclamer: I work for them.
Can someone who has this active shed some light?
Errata: I'll need a research team and a year and a half.