Google’s self-training AI turns coders into machine-learning masters
technologyreview.com
technologyreview.com
Problem is, when Google says "AI" they mean deep learning, on ginormous datasets with humongous clusters of GPUs. That don't scale.
Accordingly, when Google says "we need to scale AI out to more people" what they really mean is "we need to make more people use our services".
Sure- but most developers would be most happy with an AI system with the same predictive power as deep nets (or thereabouts) that used let's say 1/1000th of the data and could be trained on a cheap laptop in a couple of minutes.
That capability means either abandoning deep learning as inherently impossible to train small, or spending major resources to make it possible to train deep learning on small datasets with little compute.
Is Google working on this problem? Is anyone?
Try yourself at https://nanonets.com/
Of course there's the opportunity for transfer learning where you try to adjust a pretrained model slightly for a different task, but that only works if the model already mostly solves your problem. (e.g. it picks up on the necessary features, but the class you want is not one of the possible outputs.)
If your data is unlike any popular task with large datasets available, (say because you're trying to predict something from a small in-house database), transfer learning can't work (because there's no knowledge you could transfer) and you are better off with simpler models.
yorwba below explains why transfer learning is not a way to train a deep net with small data sets. Rather, it's a way to re-train an already trained deep net. If you don't have a deep net trained on your target domain, then what?
Software devs are used to having access to libraries that make our lives simple. You don't have to invent a new language and write a compiler for it everytime you want to code, because java etc. That is the hope of services like the one discussed above.
However, because trained deep nets are very task-specific and it's virtually impossible for a small developer team with few resources to train one from scratch to reach state-of-the-art performance, developers who choose to use pre-trained deep nets as a library are perpetually dependent on Google, or whatever big company decides to offer its deep nets as a service, to determine the range of uses of their re-trained nets.
...it's not really a small dataset then, though.
there are a large amount of applications where you can use big models trained on generic tasks where the data is free and available to everyone, and then fine tune to your specific more limited data.
let's say I need a person detector for my cctv camera. I use a pretrained ResNet ImageNet model distributed by Google, that was trained on a super large dataset identifying things like animals, furniture and whatnot on a huge free dataset, and then train it just a bit more on my specific task.
this allows to create very performant networks with very little task specific data.
E.g. there are so many scenarios where even really basic statistical approaches like bayesian models can provide drastic improvements over what people tend to do, but most developers I've worked with don't know how to take advantage of even that and/or don't know when they can use them or how to communicate to stakeholders what capabilities are available.
There will be ridiculously many low hanging fruits in that area for years to come.
But they had a very simple metric right in front of their nose:
Their "product" was bookings at their customers restaurants. They were paid per cover.
One trivial improvement over the constant manual tweaks would be to simply record the probability of a user choosing to book at a given restaurant if that restaurant is present in the search result for a given term or set of terms.
Once you have that data, you can use Bayes theorem to take a set of tokens for a search and produce a list of probabilities that a given restaurant will be a good match, and rank accordingly. And you don't even need to understand Bayes theorem (though as someone who usually don't like maths papers, Bayes paper is remarkably understandable with very little mathematical background) - you just need to be aware of it and be able to use a simple library.
[you'll want to do some tweaks to account for the difference that position in the result makes]
There'll be an endless stream of improvements and more advanced methods you can use, but the beauty of applying Bayes theorem to data like this is that it is very simple, and you can explain to people why it ends up ranking the way it does very easily (the action you record, such as the decision to make a booking, serves as an "upvote" effectively).
There are lots of more advanced approaches you can take if you have the time and the skill and sufficient data, but the above works reasonably well with very little effort and is very often a substantial improvement over whatever manual attempts people make at guessing how the ranking ought to be. And it can be applied not just to search results, but to pretty much anything where you can derive an input set of terms (be it e.g. categories, or a paragraph from the text on the page) that has some likelihood of having some relation to how people will like another item, and where you can record an action against a related item as a "vote" for that connection.
(I have no affiliation to Gamalon)
> Is Google working on this problem? Is anyone?
Yes, but this is very much out of fashion at the moment (just as NNs were out of fashion for about three decades). Look at the work of Gary Marcus at NYU or Michael Whitbrock at IBM.
(We are also, and think we're farther along than anyone else but aren't talking about it publicly yet so no need to believe me. There are a few others in a similar state)
Looks like Baidu's new voice cloning tech is able to do tricks like make a British man sound like an American woman with just a few seconds worth of audio to train on. Apparently the previous version needed at least half an hour of audio to train.
https://thenextweb.com/artificial-intelligence/2018/02/26/ba...
You have every ML vendor (disclaimer: My competitors) following google with this name trying to get a piece of the hype pie spreading more confusion in their marketing material. It drives me nuts.
You have folks doing everything from claiming hyper param search is "automl" to transfer learning + grid search + "insert random architecture search" here is magic that will save us all from needing to understand how this stuff works. People it takes more than that. Please read the papers for yourself down below and try to understand the limitations of these techniques.
Granted, it's great that we are attempting this, it's a real step forward, but please call it for what it is.
Here's the papers referenced in the original blog post(https://www.blog.google/topics/google-cloud/cloud-automl-mak...) :
Learning Transferable Architectures for Scalable Image Recognition, Barret Zoph, Vijay Vasudevan, Jonathon Shlens, Quoc V. Le. Arxiv, 2017.
Progressive Neural Architecture Search, Chenxi Liu, Barret Zoph, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, Kevin Murphy, Arxiv, 2017.
Large-Scale Evolution of Image Classifiers, Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Quoc Le, Alex Kurakin. International Conference on Machine Learning, 2017.
Neural Architecture Search with Reinforcement Learning, Barret Zoph, Quoc V. Le. International Conference on Learning Representations, 2017.
Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning, Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, Alex Alemi. AAAI, 2017.
Bayesian Optimization for a Better Dessert, Benjamin Solnik, Daniel Golovin, Greg Kochanski, John Elliot Karro, Subhodeep Moitra, D. Sculley. NIPS, Workshop on Bayesian Optimization, 2017.
> Automating the training of machine-learning systems could make AI much more accessible.