Cloud Video Intelligence API
cloud.google.com
cloud.google.com
Look at their example:
Animal: 97.76%
Tiger: 90.11%
Terrestrial animal: 68.17%
So we are 90% sure it is a tiger but only 68% sure it is a land animal? I don't think that makes sense.It could be that this is a weakness of seeding AI data with human inputs. I can believe that 90% of people who saw the video would agree that it is a tiger, while fewer would agree it is a terrestrial animal, because they don't know what terrestrial means.
Neural network outputs are not probabilities. I think that's the main lesson here.
The issue at hand is training them on overlapping classes :-)
I will shut up now, sorry for nitpicking.
But, I'm pretty sure the assumptions of logistic regression are even stronger than just that. The inputs are assumed to be independent given the output class, and the log odds of the output vary as a linear function of each input. The first one is essentially the naive Bayes assumption, and the second one is completely unreasonable for almost any problem ever (roughly equivalent to assuming every dataset has a multivariate normal distribution). If they are both correct, though, you get a perfectly good Bayesian posterior probability of each output class.
I think the lesson is that gradient descent will build a decent function approximation out of pretty much anything powerful enough, which is why neural networks still work even when probability theory has been thrown completely out the window.
A visual classifier that identify "4 moving things" might indicate some kind of land animal, or slow motion video of a Dragonfly in flight.
Sample/"evidence"-based reasoning will always have these kind of odd inconsistencies - I'm not sure if mapping such output to a logic model is an improvement. It might be - to take output from a classifier like this, and plug it into an expert system like a Prolog/datalog database or something. Or it might just end up being just as limited as those systems already are.
But when one says "tiger implies terrestrial mammal (or animal)", one is really talking about ontologies -- perhaps training the classifier to come up with things like "90% sure four legs, 60% sure fur" and plugging that into a logic based system would yield good hybrid systems?
I do think one would then loose the "magic" effectiveness of these pure(ish) learning systems though? Perhaps someone more familiar with the domains might shed some light?
Could be a carpet with the image of a tiger, or a toy, or a person in a costume.
For example, can the 'stripe model' or 'fang model' be shared among tigers and carnivores etc.
I tried something similar for my thesis, but this was before the advent of DNNs etc.
I imagine that they have something similar in house that they run since it is pretty vital to their core business, but you never know.
For volume consumers, Google should be paying Snapchat to learn from their data.
1: http://www.recode.net/2017/2/2/14492026/snap-ipo-2-billion-c...
Embedding Watermarks into Deep Neural Networks
Any photo you save in the app is already categorized by content.
The tipping point on training+pay for a human vs an ML/CV system for a lot of tasks is coming down.
I saw a Microsoft spotlight last fall on a company that's using drones + CV to inspect power infrastructure in Scandinavia. As much as the tech to do that has cost, its probably cheaper than sending out helicopters with specialists all the time to check power line towers for wear/damage.
Right now, most of these efforts work at the same speed humans work (probably because we still need to monitor them to make sure they're doing the right thing). But imagine self-driving cars with no traffic lights, because the cars can react & communicate fast enough to avoid a collision without any macro-scale signaling. Manufacturing processes where molten metal is shaped directly into whatever shape desired, where computers do real-time calculations of the effect of gravity & cooling rate on the position of the material. Airliners that never land and can take on passengers anywhere, because passengers ascend in a personal drone that mates with the mothership under computer control.
Unfortunately CVaaS isn't really suitable for this, because the network delays transmitting the data to a datacenter will outweigh any speed benefits of computerizing it. But I could see a big market for services that let you train a model in the cloud at high-speed, and then download the computed model for execution on a local GPU or TPU.
Once you have real time local inferencre on your mobile phone. Boom!
I'm pretty excited about precision agriculture, but for plant-life on earth it will mean that now really there will grow nothing outside the system. Bots are going to monitor all plant life. It might be plant-utopia for some species, with timely water and nutrient dispensing, but for other species it might mean being automatically killed by agrobots.
It's really scary to think about the long-term consequences of this. You know what evolution does to features not required for survival anymore.
Also if you send me an email at bookman@google.com I'll be sure to update you as to when the styling errors are resolved.
(Disclaimer, I work for google cloud)
Also, see my previous comment on a similar issue with Google Cloud Calculator:
Sorry for the inconvenience, and please don't hesitate to forward any other bugs you find to me at bookman@google.com
High-quality custom model training as a service seems much more compelling.
Of course taking as an example for a sport’s clothing company where their digital assets are mostly related to their products a general purpose API might not get the subtle differences between two similar shoes, or clothing lines from different seasons. But it might be enough to help catalog that one image or video has a sports person in, and the other is a fashion shoot or product shot.
Disclaimer: I work on Google Cloud, but not Vision.
That means you might have an interest in sailing, travel offers, or outdoor equipment, and Facebook can test that theory, and see if related advertisements have higher conversion rates.
It's also beneficial for user engagement, and by analyzing your photos, Facebook could recommend related groups in your area, or upcoming sailing events.
On the other end, imagine you make cosplay outfits for a living. You want to promote your business. Wouldn't it be efficient if you could only show your advertisement to Instagram or Snapchat users that post a high percentage of cosplay photos? That would result in a much higher click-through rate, and they could charge a higher premium for those targeted advertisements.
Or, what if you run a wallpaper site, where users upload wallpapers? You could automatically categorize and tag those images for users to search. Or, if someone is viewing a wallpaper of a sunset, you might want to show related sunset wallpapers that could interest them. That's pretty powerful. You could upload one million photos, and with a little work have them all nicely arranged in categories.
Custom computer vision models are good if you're looking for depth in certain categories. For example, if you're a gardening app and you want to take a pic of a flower and be able to recognize different species of flowers, then custom training is required.
There are some options with computer vision API companies where they will let you do custom model training. IBM will do custom training as a service for $$$$ but if you don't want to pay like crazy, Clarifai has a free (to a certain point) offering that lets you train a custom image recognition model on your own https://developer.clarifai.com/guide/train#train
People post millions of hours of themselves in natural settings on Snapchat. If you can recognize their settings (objects, environments) and cluster/categorize users then you can target advertising even more intrusively than Google et al already do.
Getting proper structured metadata from content has traditionally been expensive as it has required humans, sometimes trained as librarians, so providing the ability to extract some meaning from video becomes valuable.
Even for systems that have trained librarians, it can still help to have the ability to have a system that highlights the general content so they can further refine it and bring it inline with a taxonomy.
Google Vision API also helps because the metadata can be placed against timecode, further helping the ability to search for particular subjects within a video that might otherwise not be found easily from quickly browsing the videos.
There are other use cases depending on the customers need (such as a customer making sure that certain things are not in video about to go to air), but none of it has to do with marketing.
It's just a relative measure of confidence, scaled such that they all sum to 1.0.
Please someone correct me if I'm wrong, but I'm pretty sure that's how it works, just like how logistic regression gives you a probability.
I am working on Deep Video Analytics an Open Source Visual Search and Analytics platform for images and videos. The goal of Deep Video analytics is to become a quickly customizable platform for developing visual & video analytics applications, while benefiting from seamless integration with state or the art models released by the vision research community. Its currently in very active development but still well tested and usable without having to write any code.
The value prop is similar to, say, Twilio. Though, arguably, it's easier to run your own pre-trained CNN than it is to replicate the telephony, VoIP, and video conferencing stuff Twilio provides.
Also, presumably Google is hoping that they can continue to train and improve their CNN so that it's always just a little better than the best free-to-download ones.
There're papers like [1] where CNN output is used as input for RNN, which performs deeper context analysis. Results aren't exciting, though.
At the very least in this release they are identifying scene changes, which is a dynamic property.
The Google implementation can also detect when different people are speaking which is useful - though it takes someone to tag who is who.
As I mentioned in another post, it can also highlight where things are happening in a video - for instance a 2hr video where a scene suddenly appears. Like in a nature video rushes where an animal appears for only a few seconds.
Other types of video analysis can also detect problems in the video/audio, such as dropped frames, noise or colour gamut issues.
Disclaimer: I work for Google but am definitely not a lawyer and can't authoritatively speak for Google here.
it would be completely naive to implement it that way, considering there is an entirely new attribute video applies over images which of course is "time".
I don't know shit about ML- talking out of my ass here- but I'd be surprised if the algorithms didn't account for changes over time or canonical entity recognition (is this the same boat that was in the last image)?
> nouns such as “dog,” “flower” or “human” or verbs such as “run,” “swim" or “fly”
that out of the way... i suspect you wouldn't need video to detect those things...
and the screenshot you're referring to is an specific application of the API... not a kitchen sink:
> It can even provide contextual understanding of when those entities appear; for example, searching for “Tiger” would find all precise shots containing tigers across a video collection in Google Cloud Storage.