Google's AutoML: Cutting Through the Hype
fast.ai
fast.ai
Here's a recording from a session titled "How to Get Started Injecting AI Into Your Applications" which illustrates case studies for AutoML: https://www.youtube.com/watch?v=O7iT1INWrqo
If you have a problem that's not amenable to existing architectures then please try to tune that, but it's an advanced problem.
The whole point of venture capital or R&D is that you don't know what will work, and expect 40 failures for 1 success. She seems to think that the way science progresses is that you iterate on a Phd thesis until it's "Done" and verified and then it becomes applied. But most Phd theses, even though that survive peer review, end up inapplicable, unused, or forgotten. Not everyone publishes a General Relativity paper. A ton of Phd papers are junk, a good number don't even have reproducible results.
The only way to know if someone will be successful is to try it and let it succeed or fail in the marketplace.
These aren't examples of just productionizing a PhD thesis - they're thoughtfully designed products that solve a real problem.
Unfortunately we've watched as many ML PhD graduates launch startups which are little more than an API wrapped around the key algorithm from their thesis. These startups nearly always fail, because they don't actually address a market need.
> The only way to know if someone will be successful is to try it and let it succeed or fail in the marketplace.
There are many ways to estimate potential market size, product-market fit, etc ahead of time. They're not perfect, but they're a lot better than nothing.
His point is that sometimes research bears fruit, and sometimes it doesn't. All universities (in my country) now have metrics for how much research must successfully bear commercial-fruit or else they lose funding from the government and the EU. As such, they have commercialisation pipelines that funnel viable commercial research into products. However, even these funding bodies don't expect that anything more than a fraction of research will be commercialisable. That isn't the point of research.
> we've watched as many ML PhD graduates launch startups which are little more than an API wrapped around the key algorithm from their thesis. These startups nearly always fail, because they don't actually address a market need.
Why do you care what someone does with their PhD research anyway?
It’s nice but … I still have to scroll about 15 pages into “cat” to see a picture which isn’t my dog. I’m receptive to the argument that this is more an advertising/ positioning move than a major advance.
Disclosure: I do work for Google Cloud, but all I'm here for is to see if that dog does look like a cat.
I'd bet it trained on something like the fur texture since it also matched things like a lemur which have longer smooth fur.
Dog vs Cat is one of the best studied problems in deep learning, and there is lots of training data. This is very surprising!
I'd bet it trained on something like the fur texture since it also matched things like a lemur which have longer smooth fur.
The line on this forum is that Google's product is its customers' data, but that's never been right. Google's product is, and has always been, dirt-cheap computing. They have a really large amount of computers, they are building more right now, and they want you to use them. The surplus of computer power within Google is what makes Googlers sit around thinking "sure, that was more CPU time than anyone has ever used for anything before, but what if we used 100x that much?"
https://static.googleusercontent.com/media/research.google.c...
"Research to make deep learning easier to use has a huge impact, making it faster and simpler to train better networks. Examples of exciting discoveries that have now become standard practice are:
* Dropout allows training on smaller datasets without over-fitting.
* Batch normalization allows for faster training.
* Rectified linear units help avoid gradient explosions.
Newer research to improve ease of use includes:
* The learning rate finder makes the training process more robust.
* Super convergence speeds up training, requiring fewer computational resources.
* “Custom heads” for existing architectures (e.g. modifying ResNet, which was initially designed for classification, so that it can be used to find bounding boxes or perform style transfer) allow for easier architecture reuse across a range of problems.
None of the above discoveries involve bare-metal power; instead, all of them were creative ideas of ways to do things differently."
These seem quite fundamental (except for the last one which sounds like a favorite domain adaptation strategy from the author).
I'm in industry myself, but it's quite hard to come up with fundamental strategies. In particular I'm trying to merge nonparametric Bayesian models with deep networks. By enriching the latent variables in autoencoders to richer priors we might see improvements. Subsequently, we simultaneously need better control variates to do inference in more complex models. See my blog post on the overlooked topic of control variates: https://www.annevanrossum.com/blog/2018/05/26/random-gradien.... If we really want to be creative we need people from academia on board.
Explainability is important, and critical in applications where lives are on the line.
Explainability where we're guessing what's happening in a black box model won't do either. Nothing but complete transparency of the model and why it's doing what it's doing. Its source code, that makes sense to humans, is needed. Full on model audit. No guessing.
I can think of only one company that's attempting to do this, and it's not anyone you hear working on explainability, including DARPA.
2. This rant has very little to do with the article, and feels like a borderline meme that some people post on all that is ML-related.
Yet we don't demand that other humans explain how their visual cortex work. There is a double standard here.
State models (Markovian) are sometimes able to explain things but not always really, especially in complex cases.
[0] https://www.csail.mit.edu/research/interpretability-complex-...
First of all: Fast.ai are a non-profit. There's no ulterior motive here. I think a lot of the commenters here are feeling clever "looking for an angle", when there honestly is none.
Secondly, I really can't think who should have accrued more benefit-of-the-doubt than Rachel Thomas. It's just silly to take the point of view expressed here as insincere or motivated reasoning. None of this means you have to agree with the points raised, the predictions made, or the conclusions drawn. Of course. But the snide tone of many of the comments here is really discordant.
Finally, it's a little...revealing, that there's so much discussion of Jeremy here, including comments that seem to assume he wrote the article. I don't even know what to say about that.
2) "We can remove the biggest obstacles to using deep learning ... by making deep learning easier to use"
3) "Research to make deep learning easier to use has a huge impact"
Because this article rants against academics, but is written by an academic, it reads as inconsistent and disingenuous. Wild guess...AutoML and Google are worrying competitors to whatever their business is?
My bet is that most of data engineering/data science/designing neural network architectures will be commoditized in 5 to 10 years. Maybe sooner.
Excuse the snark, but:
"In evaluating FastAI's claims, it’s valuable to keep in mind FastAI has a vested financial interest in convincing us that the key to effective use of deep learning is more machine learning experts, because ML education is their business model. If true, we may all need to purchase FastAI courses. On its own, this doesn’t mean that FastAI’s claims are false, but it’s good be aware of what financial motivations could underlie their statements."
I can think of thousands of ways that Google could increase the computational power required that would be much easier than the AutoML effort (for ex. simply recommending ridiculously deep models that take 100x to train without giving higher performance). They are putting so much effort into AutoML because it just works. A lot of the things included in later parts of this series are very useful (learning rate search, etc.) and decrease the computational power required, but most people just want to drop a dataset and pick the type of model (say, multiclass image classification) and leave the rest for machines to optimize.
It would be much more valid if they were a business, rather than an open source software and free course solution. Even as a competitor, it's probably useful criticism to get a negative perspective from a competitor.
I think the disconnect here is that you can reuse existing architectures and get state-of-the-art performance without running something like AutoML. It's not clear that creating a bespoke architecture for your specific problem is always better, let alone always a good use of your resources.
I've built models that worked fairly OK, only to have a colleague build a separate model that had +10% prAUC by virtue of adding some additional mechanism (say, attention, a different RNN cell, more units, etc.).
I'll also say that this is aimed towards people that are unfamiliar with ML and would have trouble finding and re-implementing the state-of-the-art in their specific field.
Arguably, my point was that this argument, biasing your analysis by hidden motives that cannot be disproven, is invalid.
But clearly, I was wrong because fast.ai's courses are indeed free so the comparison doesn't stand.
I'll also add that fast.ai has great ML courses that go deep into details, and focus on practical, applicable, techniques. My comment was specifically about the argument that Google is doing this just to increase Cloud revenues.
Data Scientists in many companies have become close to what SQL Job or any web script job runners used to do and then CRON jobs came into pic and that's what Auto ML will do here with a a set of sequential steps replacing mundance ML activities.
It just works, provided your dataset can be solved well enough given AutoML's pretrained model. Which is not at all guaranteed.
/s
Rather, it seems to me that their policy so far amounts to an attempt at platform lock-in. If you want to do Deep Learning like Google does, you basically have to use their tools (yes, that's tensorflow I'm talking about) and their data (in the form of pre-trained models) and eventually pay them for CPU or GPU time (you can also pay their competitors, of course).
In fact, I'd take this a step further and say that the whole Deep Learning hype is starting to sound like a total marketing trick, just to get more people to use Google's stuff, in the vain hopes of achieving their performance. However, to beat Google at their own game, with their own tools and their own data, running their models on their own computers... that does not sound like a winning proposition.
Poor folks must live a constant emotional rollercoaster reading this site.