Deep Learning's Diminishing Returns
spectrum.ieee.org
spectrum.ieee.org
This is one of the key misunderstandings are still deeply rooted in people's minds. For modern DL, a large part of the learning comes from "internal" data points, in this case the pixels of the image, as opposed to the labels. If you count the number of pixels, you will likely get something like 1.2 trillion, more than enough to justify the 4.8e8 parameters. It's the usage of internal data that prevents overfitting, NOT the random initialization and SGD as claimed in the article.
Another way to see this is: if you need more labels than parameters, how can GPT3 have ANY parameters at all? It is trained purely on raw text data.
Input: <previous words of article>
Label: <next word>
Your point is well taken that the number of input data points is also important when considering the complexity of the problem. In this case however the number of data points more or less exactly equals the number of labels.
(About Me: the first year+ of my PhD was focused on large scale language modelling, during which transformers came out.)
Which all adds to the number of constraints / equations.
GPT is trained to predict the input (estimating p(x)), versus predicting a label given an input (p(y|x)). So in the case of GPT you can use the input dimensionality as a "label", as another responder has mentioned. ImageNet classification is different (excepting recent semi-supervised or unsupervised approaches to image recognition).
The ability to generalize in the typical imagenet setting is, as the article says, a byproduct of SGD with early stopping, which in practice limits the number of functions a deep neural network can express (something not considered in an analysis which only considers parameter count).
Input dimensionality is absolutely important when determining net size.
Maybe I understood that question wrong, but regardless, even if early stopping wasn't implemented, a NN would have more predictive power than the hash mapping. Both would be completely overfit on the training data set, yet the NN would most likely be able to make some okay guesses with OOD data.
What you should look at is the number of outputs times the number of data points for each output. If this number is lower than the number of parameters then it should be possible to find multiple solutions.
Of course in this case you're not looking for a solution, but an optimum, and not even a global one, so it's not too troubling per se that you don't get a unique answer. Though it does somewhat suggest you should be able to get an equivalent fit with far fewer parameters, but finding it could be quite tricky.
All journals have good and bad papers, the same goes for conferences. So to understand the quality of a piece of work you need to at the very least talk to a few people in the field that have engaged with the work to form a reasonable opinion about its quality.
> ... A similar sentiment was expressed by Marvin Minsky: "Unfortunately, the strategies most popular among AI researchers in the 1980s have come to a dead end," said Minsky. So-called “expert systems,” which emulated human expertise within tightly defined subject areas like law and medicine, could match users’ queries to relevant diagnoses, papers and abstracts, yet they could not learn concepts that most children know by the time they are 3 years old. “For each different kind of problem,” said Minsky, “the construction of expert systems had to start all over again, because they didn’t accumulate common-sense knowledge.” Only one researcher has committed himself to the colossal task of building a comprehensive common-sense reasoning system, according to Minsky. Douglas Lenat, through his Cyc project, has directed the line-by-line entry of more than 1 million rules into a commonsense knowledge base."
will sublime conclusions made by GPT-3 EVER be trusted if the reasoning to its conclusion is not understood by a human? perhaps the gestalt of gpt3 implies a meta gpt3 that could derive human grokkable explains of its "dumber" self. or maybe not.
Just a hint - each of these "facts" or "rules" mean slightly something else depending on the context they are used in. This context is not hard-encodable because it slightly mutates with every new information learned or forgotten.
I'm yet to see GPT-3 do anything commercially important? Cyc on the other hand seems to have been used in a number of sectors. Not to downplay GPT-3 - it's cool tech that produces cool demos - Cyc just seems more like a tool rather than a toy.
To name a few:
- GitHub CoPilot (transformative imo)
- Markcopy.ai, jenni.ai, etc.... Tons of content generation and SEO tools startups
- AI Dungeon and such
- Plenty of chatbots
- It's super useful for all kinds of classification tasks too (as are all transformer models)
edit: Cyc can also tell you why it gave the response it did - something that deep nets cannot do. This is important in many fields, otherwise you cannot trust any response it produces.
Contact me if you want to see it applied to text based DL models, time series based DL models or Image based DL models.
The following argument comes to mind, but I don’t really buy it (it just came to mind as something that one might say next):
‘ Perhaps there is an analogy between the solutions of “we just need to get better (more varied, better fitting the desired behavior, etc.) training data, and maybe better training procedures” and “we just need to add more/better inference rules and symbolic ways to encode statements, and add more facts about the world”. Similar in that both will produce the specific improvements they target, but where solving “the real/whole/big problem” that way is infeasible. If so, then maybe this indicates that a practical full-solution to artificial “common sense” would require something fundamentally different than both of them, if it is even possible at all. ‘
Again, I don’t really buy that line of reasoning, just expressing my inner GPT2 I guess, haha.
Ok, but I presented an argument (or something like an argument) which I made up, and said that I don’t buy it. So, I should say why I don’t buy it, right? Like many of the things I write, it is chock-full of qualifiers like “perhaps” and “maybe”, to the point that one might say that it hardly makes any claims at all. But ignoring that part of it, one major difference is that the DL style architectures, seem to be working? And it isn’t clear what kinds of (practically speaking) hard limits it could run into. Now, on the other hand, perhaps at the time that symbolic AI was all the rage, it appeared the same way. (Is this what people mean when they talk about inside view vs outside view?).
Why should these two things not be especially analogous? Well, saying “proposed solution X to the problem says to just [do more of what X is/do X better], and that is just like how proposed solution Y says to just [do more of what Y is/do Y better]” is kind of a fully generalize argument for dismissing any proposed type of solution where partial solutions of that type have been tried, but the whole problem hasn’t been solved that way yet, and another proposed kind of solution has already lost favor. This doesn’t seem like a generally valid line of reasoning. Sometimes you really do just need more dakka (spelling? I mean “more of the thing you already tried some of”).
Of course, if one is convinced that it really was right for the older proposed kind of solution to be discarded, that probably should say something about the currently popular kind of solution. Especially if there have been many proposed kinds of solutions which have been discarded. But, it seems like much of what it says is just that the problem is hard. And, sure, that may mean an increased probability that the currently popular proposed kind of solution also doesn’t end up being satisfactory, that doesn’t mean one should be too quick to discard it. Tautologically: if no known alternative is currently at least as promising as the type of solution currently being considered, then, the current one is the most promising of the currently known options. Whether it is promising enough to actively pursue may be a different question, but it shouldn’t be marked as discarded until something else (perhaps something previously discarded, or something novel) becomes more promising.
That said, I don't think we are going to be able to actually step away from large networks any time soon. It seems to be the case that when you have more parameters to optimize, you have more degrees of freedom and you are less likely to end up getting stuck in a local minima which is why it's actually easier to train a larger network to solve a task versus a smaller network, despite the fact that both are wildly overparameterized.
Is this true for all deep learning models?
That's why the current Tesla has like Five LIDARs, two radars, 16 cameras, and lots of traditional algorithms that aren't NNets. (I don't know the counts, but my point stands.)
I've spent 5 years on this with tier 1 automotives, I don't think ADAS5 will happen in my lifetime, and it definitely won't be due to neural nets: It is FAR more likely autonomy will be achieved by billions of IoT sensors in a mesh, guiding vehicles: e.g., lane sensors, weather sensors, c2c (car-to-car telemetry), c2i (car-to-infrastructure telemetry, like traffic lights), rather than some fancy in-car brain-emulation.
Some of my favorite tesla fails:
[1] https://techstory.in/tesla-autopilot-tricked-by-yellow-moon-...
[2] https://jalopnik.com/tesla-autopilot-glitch-makes-for-someth...
[3] https://www.theregister.com/2019/04/02/tencent_tesla_hacking...
There's a good reason that consumer software either has small n, or basically doesn't use algorithms with exponents worse than n^2 (and, we only go to n^2 if it's really necessary). Throwing compute resources at the problem only goes so far when that exponent works against you.
At that time, I wasn't thinking about it in terms of energy consumption, or carbon emissions, but, given the scale of things now, that does look like an appropriate set of units. Nor did I ever conceive of algorithms that would only be run once, because it was just too expensive to run it more than that.
The sum of all DL training in teh world is noise compared to the other big consumers of energy in computing. That's because the main players all invested in energy-efficient architectures. DL training energy is not something to optimize if your goal is to have a measurable impact on total power consumption.
If the cost was gigantic enough to make the investment worth it they must have found some really great improvements for it to end up being just noise. Improvements that somehow didn't have a noteworthy impact on general computing.
A big company like Microsoft probably wasted more money on pentium 4s 15 years ago. Electricity is just another resource - if the numbers work, burn away.
I know we’re nowhere near the following scenario, this is just to illustrate how things can go wrong even if the numbers tell you to “burn away”:
Image we have computronium with negligible manufacture cost, the only important thing is the power cost to use it.
Imagine you’re using it to run an uploaded mind, spending $35,805/year on energy.
The 50% of Americans earning more than this [0] are no longer economically viable, because their productivity can now be done at the same cost by a computer program.
Doing this with the current power mixture would be disastrous, doing it with PV needs about 1400m^2 per simultaneous real time mind upload instance (depending on your assumption about energy costs and cell efficiency, naturally).
In a more near-term sense, there are plenty of examples where the Nash equilibrium tells each of us to benefit ourselves at the expense of all of us. Not saying that is the case for Deep Learning right now, but can (and frequently does) happen.
I hate to be the one to tell you, but, it turns out we are living in the middle of an ecological catastrophe, and it also turns out that means that electricity is a resource we are going to have to conserve.
Which is many orders of magnitude more energy-intensive, on the scale of a small nation-state, and in most cases fundamentally wasteful by design. A very large pre-trained model can be reused very cheaply once it's finished.
> Also from my limited knowledge I think DNN models are not very transferable in real world setting requiring constant retraining even for a small drift in signal or change in noise modes.
This is FUD, promulgated by people who expected deep learning to solve all their problems overnight. All models will suffer from "drift" whenever the underlying data changes.
Part of what made deep learning so good was that it was able to generalize exceptionally well from exceptionally complicated input data.
It is unreasonable to expect that a model pre-trained on a huge generic corpus will be a perfect match for your very specific business problem. However it is _not_ unreasonable to expect that said model will be a useful baseline and starting point for your very specific business problem.
We are not yet (and might never be) at the point where you can dump a pile of garbage data into an API and get great predictions out the other end on the first try. But nobody ever thought you could do that, except the people selling expensive subscriptions to those kinds of APIs. The fact that they work at all should be taken as evidence of how amazing deep learning is; the fact that they don't work perfectly should not be taken as evidence that deep learning is bad/useless/wasteful/hype/whatever.
Don't let the clueless tech media set your expectations.
Professional data scientists and machine learning practitioners for the most part take their work very seriously and take pride in delivering good outcomes, just like professional software engineers. If deep learning wasn't useful to that end, nobody would be using it.
Facebook trained image understanding CNN ("Lumos" in the paper) only once in many months, but this is not usual.
Those claims are entirely new to me, and I've been a researcher in the field for almost 10 years. Where do they come from/what theorems are they based on? It's unfortunate this article doesn't have any citations.
https://en.wikipedia.org/wiki/Cram%C3%A9r%E2%80%93Rao_bound
But it only applies to estimation (like how well a population parameter can be estimated) in certain regimes, and not e.g. expected risk (like how well one can do at prediction), so I’m not sure how it would apply here.
The article also ignores training vs running tradeoffs. Training a model once may be extremely resource intensive, but running the resulting model on millions of devices can be negligible while having huge value add.
a good example is discovery of attention mechanisms/transformers replacing more cumbersome and computationally expensive RNNs and LSTMs in NLP and more recently outperforming more expensive models in computer vision.
This is what’s called infeasible.