Cloud AutoML: Making AI accessible to every business
blog.google
blog.google
I think that if AI is to be accessible to every business then it will deliver insights rather than the machinery to produce the insights. This is especially true in the context of small businesses.
I would also hesitate to build a business relying on Google for those things since I'd likely be competing with Google's actual moneymaker.
Can you say more about the specific insights you'd like us to provide? The more specific the better :) Happy to see what we can do!
However, I'm curious if they plan to support structured/relational datasets which are definitely something every business needs. In Kaggle's 2017 State of Data Science [0] survey, data scientists said they spent 65% of their time using relational datasets vs 18% for images. Given that Kaggle is owned by Google, this must be something on their radar.
For those data scientists, I maintain an open source library for automated feature engineering called Featuretools (https://github.com/featuretools/featuretools). For people interested in trying it out, we have demos (https://www.featuretools.com/demos) to help you get started.
It will be interesting to see the trend over years. One year doesn't say anything about the trend in the industry.
Part of the reason is that use cases are driven by limited definitions of ML or "AI" -- AutoML for example note they are "Making AI accessible to every business..." but they are just managing a small part of one type of ML (images with conv nets and res nets.) A senior exec who reads this might develop a narrow view of what AI is.
A general annoyance is how obvious techniques like regression, tree models, Bayesian models, etc on tabular data are so ignored while everyone gets hyper-obsessed over GANs or whatever. Almost 90% of low-hanging-fruit I see can be captured with simple classic ML applied to tabular data.
We're just getting started! Stay tuned for lots more AutoML goodness.
If you compare ML to electricity, we're still in the stages where a few players have found that electrifying their manufacturing plants makes sense. Small players can't afford he investment in machinery and skills. Maybe when the machinery is hidden behind a "utility" provider (which would also bring down the skill level) they will.
Bear in mind that this service creates the architecture of the model for you. I think ClarifyAI has a predefined model that is fine tune with the data, which it is not even similar.
I wonder if Google tests for Doublethink skills before you can get hired there now.
what if in the future they change their minds and decide to change the TOS? did I just build on top of quicksand?
The data remains yours and the model is yours - eg if you delete your account, the data and model goes away (think of it like the data or model you stored on a VM). However, what I think you're looking for is to be able to actually download the model, and I'm afraid that's not possible.
Are you looking to avoid lock-in? Or something else?
Eventually useful for building a very basic model to serve as a benchmark for a real model.
There are a few different reasons:
1) Most companies have very limited skill in building advanced models. While many folks are trying to achieve 5% or better accuracy, MOST folks are trying to achieve ANY accuracy (since they have nothing)
2) Many problems are not very complicated and do not require a custom model. From the blog post, the "cloud" example only requires a small amount of changes to classify for a specific domain - to have to build an entirely new model for that, or train on MM of images seems like overkill
3) AutoML (often) is better than humans already[1]. So if you want to achieve that 5%, you MIGHT need to use a machine anyway.
And second, your comment that most folks achieve any accuracy is strange. These are not real businesses, they are mostly developers and hobbyists trying to learn. These folks sign up for kaggle and poke a few scripts and view half a class on coursera on ML. They are not real businesses and they have no money. Most of the real businesses are hiring startups or large companies that hire data scientists with domain expertise like in oil, manufacturing etc. (the IBM model). ML as an API is a disaster as a business model.
Also, AutoML is not even close to being better than humans even for a specific problem (across datasets). These click-bait titles don't fly outside of AI conferences.
Put another way, the average customer has ~zero ML usage today. I'd guess that 95%+ of all businesses have ML usage today. Further guessing would say that <1% of ML usage actually care about levels of accuracy beyond "it's better than the hacked together set of rules/filters I use today." These are very large businesses with lots of money to spend on a solution.
There are many ways to measure "better", and AutoML does apply here. This includes "better == faster to train or develop" [1], "better == you need less data" [2], "better == lower error rates"[3]. While I agree that many of measures do not apply across datasets, most customers only have one dataset per problem.
[1] Predictive accuracy and run-time - https://repositorio-aberto.up.pt/bitstream/10216/104210/2/19...
[2] Less data - https://arxiv.org/abs/1703.03400
[3] https://link.springer.com/article/10.1007/s10994-017-5687-8
Disclosure: I work at Google on Kubeflow
- Average customer has zero ML
- Nearly no customers are using any ML (difference between median and mode)
- Of those that are using, very few care about better than human perf
Is the argument that it is easy to implement stock models but hard to tune the models for specific types of image inputs? Inst that pretty easily solved with some parameter grid searches? How much specialized skill does it take to re-do networks from traditional inception architecture or what not into something specific for hot-dogs or satellite imagery or medical images?
* half the student interns I interview from top-20 comp sci programs do this on weekends for hackathons*
It's trivially easy to take what someone else has built and modify it slightly for a similar problem, especially in a hackerthon environment where you can ignore edge cases etc.
See if they can build a new model from scratch for a new type of problem. I'm not saying that AutoML can do this either, but I interview large numbers of PhDs who don't know where to start on doing something new.
"It's very basic but you can use the model as a benchmark" is actually the pitch you get from AWS.
Correct, it's a cloud service, based on the research Google published for model exploration[1][2]. There are research examples today where this service provided better models than humans were able to achieve by hand or with genetic algorithms (models trained faster and/or with better error rates)[3]
[1] https://static.googleusercontent.com/media/research.google.c...
[2] https://blog.acolyer.org/2017/10/02/google-vizier-a-service-...
Note I left it off of this one. :)
I'll make sure we have clear statements/auditing on this! We're already committed to GDPR[1] - is there a specific additional service/certification you'd like to see?
If Google is worried about people just generating models and running, then charge for training time/resources.
That's fair - the reality is that it's less about the worry of people "taking of models and running it themselves", it's more about the amount of custom stuff that has to go on behind the scenes in order to enable this kind of model exploration and training. We're exploring all options but, yes, you'd be committing to the platform (for now).
If you're interested in a portable solution, may I recommend (extremely selfishly) Kubeflow[1]? We don't provide AutoML, but we do provide a very portable solution for running an ML stack.
Well.... no. BUT, it _does_ support unstructured data so you may not need to label your data at all. As always, YMMV.
Integration with human labeling
For customers with images but no labels yet, we provide a team of in-house human labelers that will review your custom instructions and classify your images accordingly. You will get training data with the same quality and throughput Google gets for its own products, while your data remains private. You can use the human labeled data seamlessly to train a custom model.
It's great to see that other cloud providers are acknowledging the talent and training data gaps that many large enterprises face when adopting deep learning.
Disclaimer: I work for AWS
This is an externalization of the service we use at Google internally called Vizier[1], first discussed publicly in June[2].
The idea is that instead of having to build a model yourself, we can use ML (yes, it uses ML to provide ML) to autotune your model and solve your business problem. Basically, instead of having to deal with all the steps in opening an editor, choosing a algo, tweaking, debugging, etc etc, just provide your structured or unstructured data and we'll help you answer your question (which is what customers actually care about).
[1] https://research.google.com/pubs/pub46180.html [2] https://www.youtube.com/watch?v=Z2YL4XJKVpQ
This significantly reduces the level of expertise required to train models, and the AutoML models outperform "expert" human-created architectures.
Interesting! I read up on Sagemaker here[1] and didn't see any AutoML style training/tuning features, but you would certainly know better than me :)
(I know you are right, but so many people here think AutoML is exactly the same that the HPO they were doing since long time ago)
AutoML vision is actually built on Google Brain’s proprietary image recognition technology, and Vizier is one of the components of their broader solution. You can see their earlier research announcement here[1]. Sorry to leave off the additional teams that helped in building this!
[1] https://research.googleblog.com/2017/05/using-machine-learni...
Many businesses need them, but don't have the staff or expertise, many may want them for fun functionality (building their own "Not Hotdog" app), but the aim is ease of model creation for anyone with a bit of data and time.
(Disclosure: I work in Google Cloud)
It's a start!
(Disclaimer: I work for Google but not on this.)
The post says there is transfer learning involved, which means in practice you need much less data than you would if creating a classifier from scratch. Of course, more (good) data may yield better results, but it seems one of the goals behind this release is specifically to give custom (your own labels, not just generic object detection) high performance image classification to those who don't have access to Google-scale training sets.
Can you say more? I don't think anyone is saying it's magic pixie dust, but it does dramatically reduce the amount of data you need.
I don't doubt Google has managed to make something useful work, though I'm more skeptical of how general the ML tech is. One advantage of an API like this is that it allows control over many of those variables. I'm not sure if this is what it does, but you could even start out by making a transfer-learning system that's heavily tailored to transfering from one specific fixed model, which coupled with some Google-level engineering/testing resources, could produce much more reliable performance than in the general case.
I've also seen it used to justify insufficient validation - resulting in strange generalization failures.
As you can see here[1], we do provide quite a bit of information about the accuracy and training of the underlying model.
Additionally, the AutoML already (often) provides better than human level performance[2]. Your comment about transferring a heavily tailored model from one model to another is basically what it's doing - it's taking something domain specific (vision) and allowing you to transfer it to your domain.
[1] https://youtu.be/GbLQE2C181U?t=1m15s
[2] https://static.googleusercontent.com/media/research.google.c...
Besides, the fact that you said that linear regressions can be useful for small datasets is exactly my point.
Source?