Deep Learning’s Impact on Image Processing, Mathematics, and Humanity
sinews.siam.org
sinews.siam.org
Examples of elegance in deep learning: The gating mechanisms of LSTMs (also adds to interpretability); structure of convolutions, i.e. weight sharing, translation invariance; the policy gradient theorem (Sutton 1999); dropout, its relationship to biology and variational inference; Generative Adversarial Networks, and connections to game theory, split-brain; the reparametrization trick; the log-derivative trick; connectionist temporal classification. These examples are only the surface. Depending on your specialization inside deep learning, you'd find many more.
Examples of interpretability: Nguyen et al.'s Synthesizing preferred inputs to hidden neurons, Zeiler et al. convnet visualization, Guillaume Alain's linear probes of hidden layers, attention readouts in attentional models for machine translation or speech synthesis, and many more. Ultimately, deep learning methods are probabilistic, and decent deep learning engineers would be able to tell you why a model is doing what it's doing by printing probabilities and activation statistics, much like any other probabilistic machine learning model.
(I'm not Elad or affiliated to him or his instutition btw).
Lots of people have "experience of machine learning" these days. And there's so much information out there anyway that you really don't need to know what you're talking about to leave a quick comment on a discussion board.
It's about the application. If you are detecting cat images on a social network, then a few false positives or negatives are totally no big deal. If your algo is making investment decisions or is being used as evidence in court then it is entirely reasonable to expect it to have a rock solid causal chain easily comprehensible by a layperson. The example in the article is noise reduction and the method was "we tweaked it until we got the results we wanted". In many fields that's called "overfitting" (or "fitting up") and is a big red flag.
Citing certain examples of misuse of deep learning by folks who "tweaked it till it worked" doesn't say anything about deep learning at large. The same team would do the same with a much less powerful model.
Of course, being able to explain is a very valuable trait. But there's just no evidence that deep learning methods are any less amenable to interpretation than SVMs or decision trees. The latter models' "interpretation" is mostly stuff like "feature 543 and 632 were on" , while deep learning methods can not only do that, but can also synthesize examples of characteristics that the model looked for [2].
[1]: https://arxiv.org/abs/1611.03530 [2]: https://arxiv.org/abs/1605.09304
Yes and no. Let's say you have a time series and you know there is some pattern in it and you are interested in predicting what it will be at time t. A Fourier analysis would give you a formula to plug numbers into that anyone can understand. NN or DL might give you as good, or accounting for noise, even a slightly better prediction when backtested but now you need to explain "why did it do this" - to a financial regulator who is looking at insider trades, or a commission looking into why a piece of equipment failed killing the user (say, a self-driving car), or a newspaper who's noticed more people of race X are being turned down for mortgages. That's why I say it is dependent on the application of the algo.
SVMs are well-known black-box optimisers of numerical parameters so not more or less interpretable than neural nets, but decision tree learners are a different matter. The model they build is symbolic- a propositional logic theorem that can be output in a format that is directly inspectable and interpretable by a human being with an understanding of propositional logic.
It's true that once decision trees grow beyond a certain size they're hard to interpret, but even so they can make a lot more sense than a heap of numbers in a large graph.
Neural nets used for image processing are special in that they can output their activations as images that are directly interpretable by humans, but that's about the only application where we can really tell what a deep network is learning.
Edit:
>> All machine learning models are capable of overfitting.
Aye, but deep nets by definition build extremely complex models with millions of parameters and so are especially prone to overfitting. Bias-vs-variance and all that.
Edit: I should add that I worked on a similar problem. I can't reveal all the details, but it was about scoring pairs of objects to say how much one object is like the other. At first we used an off-the-shelf system that only spit out abstract values, so you can say "A is like B with a value of 5, but it's like C with a value of 5.17, so A is more like C than like B". The problem was those scores didn't correspond to any explainable metrics, so we switched to a custom system that could spit out a normalized value (0 to 1) and tell you exactly which qualities were alike and how the score was formed. People were NEVER happy with the values. Never. Everyone "felt" that these two should be more alike, or there should always be at least one object that has at least 0.8 likeness to another, etc. Plus, the values were now non-transitive, i.e. A can be 0.9 like B, but B is 0.2 like A. After about a year of tweaking we switched back to a non-explainable system.
Of course, the "non-explainable" system is perfectly explainable, it's a computer program and it isn't even deep learning, it's just long equations, the explanation is just "this long equation results in this value". The semantics of that equation though are very hard to impossible to understand. Plus, I don't think we can have a semantic explanation of a deep learning system precisely because we use deep learning when we can't semantically specify an algorithmic solution to the problem.
It depends who you are explaining it to, of course, different audiences will have different standards, and will have or have access to, different kinds of experts to validate the explanation.
If the question is "why did your self-driving car hit that pedestrian?" for example, the standard would be pretty high. "It works most of the time but we don't know why it failed this time and we don't know when it might fail again" isn't going to fly. If the question is "we've just busted this trader for insider dealing, and you made the same trades at the same time, why?" then "the computer just told me to for no particular reason" isn't going to go down well either. You would need to show how data that you legitimately had access to reproducibly leads to exactly the same decision being made.
What do you do if you don't like the explanation? That's easy. Someone's going to jail. Don't let that be you :-)
It didn't see him/her. Which is the extent of an explanation a human can give you. Try to picture wanting a person to explain why they didn't see a person coming from the side. We don't know what's happening inside our own heads.
That's my point - we're recreating an inherently non-explainable system in software, because we can't explain how it works. The algorithm is so complicated, we can't code it explicitly, so we let the machine figure out the semantic-free equations. All is not bad though, you could always dry-run the algorithm at home, i.e. ask your car "how visible am I with these clothes on", and just not walk outside with low-visibility and low-recognition clothes. Some countries already require pedestrians to wear reflective clothing in the evening to make themselves more visible. I suspect we'll have similar rules for self-driving vehicles: you must wear recognisable clothing when walking through traffic.
What I mean about "what if you don't like the explanation" is something else. Let's say you use a deep-learning software to determine sentences. The software spits out "5 years, because this is a black male". Do you accept that? It has been proven effective in many previous cases, can you accept that it has independently discovered a correlation between the person being black and the length of the sentence? Let's say you entirely remove the person's colour and sex from the dataset and re-train, then it will give you shorter sentences than you expect, because every judge previously had subconciously taken all these variables into account. Or it will discover a correlation between wearing e.g. adidas and alcoholism because Slavic people have a higher rate of alcoholism and wear adidas more often than others.
Which comes back to my core argument: If you know the acceptable metrics and algorithms, why are you using deep learning? If you don't know the acceptable algorithm, how can you verify the explanation?
There is a world of difference between "this individual driver made a human error" and "software that can't reliably see people is driving a million cars now". The developers of the software would need to demonstrate they understood why the software failed in this situation and how they are going to fix it.
I suspect we'll have similar rules for self-driving vehicles: you must wear recognisable clothing when walking through traffic
Bear in mind that we can't even make cyclists now wear hi-vis or fit lights - making the entire population wear certain things for the benefit of job-destroying technology isn't going to get many votes!
If you know the acceptable metrics and algorithms, why are you using deep learning?
I disagree, because it's a matter of scale and usefulness. An algorithm that detects pictures of cats can answer the question of "why do you think that's a cat?" with "because it looks like a cat" and that's a perfectly reasonable explanation that a human might give[0]. But the stakes are raised when you start to deal with real people in the real world. "Why did you turn down that loan application" or "why did you add that name to the no-fly list" or "why did you reject that job candidate" need a bit more exposition.
That's a self-correcting problem because no-one sane will risk their business on software that "inexplicably" just happens to reject everyone in a protected class because there happens to be some vague correlation between members of that class and some undesirable activity or outcome.
[0] Except http://www.bbc.co.uk/news/technology-33347866 of course
Which would make Driverless cars a non starter in the real world.
Dear God, I can't imagine the poor coder who has to dictate fashion so that car firms can ensure they don't hit people.
That addition reneges on the basic expectations people had of driver less cars in the first place, so now why would they have any motivation to let them happen?
Let's say we ban development of self-driving cars since the algos can't be explained. Then we will never know the benefit we might reap from adopting them. We know humans make mistake. But with self-driving cars, the error rates in future might be very low as compared to human. We as a human society will never experience this future because we had this silly idealism that all algos must explain themselves in human terms.
That's a false dilemma since it's by no means proven that the algos CAN'T be explained - just that we don't know how to extract the explanation from the model yet. It's entirely reasonable to suspend real-world use until the maths catches up.
It's troubling when we don't know why an algo is racist. We need ways of checking that these influential algos aren't reinforcing trends we would rather diminish.
reinforcing trends we would rather diminish
Well, of course they are, at the end of the day all any predictive algo is doing is extrapolating a trend. It usually requires serious regulatory intervention - such as men being charged more for car insurance because "the computers said so" until legislation was passed barring sex as an input to the model.
Additionally, the study about the "racist algorithms" was fraudulent. Their results were "almost statistically significant". I.e. not statistically significant. Compared to human judges which have well studied biases by race. There's nothing remotely fair about human judgement and you should always prefer an algorithm. In almost every domain they could find, researchers have found that even simple linear regression beats predictions of human experts.
And this has huge effects on our society. Humans making biased hiring decisions leads to mass discrimination against certain groups and very suboptimal employees. Having humans make loan decisions, means much higher interest rates, more people go bankrupt, and the economy grows much slower.
The EU banning the use of algorithms is just absurd.
You assume that your opponent is reality.
your opponent is other human beings. It takes precious little for a motivated person to learn how to hide malfeasance behind a black box.
Faith in un-corrupted black boxes should be considered the same way as faith in the un-coruptible internet
The average programmer at a tech company won't be able to tell us how a particular complex piece of code works, but that doesn't stop us from building complex software.
Deep learning methods are also not off-the-shelf type algorithms. Using them does require knowledge of the domain. This doesn't fit with the "black-box" narrative.
In fact, SVMs and DTs are black-boxes due to their off-the-shelf nature. (jk lol)
Well, nobody's against using the brain, but the current trend in most subject domains is to avoid overrelying on intuition, and checking and supporting the conclusions using structured thinking methods, i.e. logic.
That is, deep learning is an intuition, and you have to have a high-level explanatory/verification mechanism that would support or reject the answer.
Take the recent discussion on reddit/ml for example, people are still debating about whether it should be conv-bn-relu or conv-relu-bn. This is a pretty widely used building block, if not the most widely used one, however, people still don't understand why the latter could work or even outperform the former in a lot applications since it filters out all negative values thus destroying/skewing the underlying distribution for bn. And for BN alone, there is a lot of questions to ask, like the running statistics feels like a hack, however it works very well in reality.
So I take no issue of calling deep learning nowadays a black box. We are far, very far from understanding why this monster does this well in solving so many problems. That is why it is interesting. Some researchers' attitude is confusing to me, because apparently there is a big juicy problem out there, waiting to be cracked, yet, they are distancing themselves away from it.I cannot help thinking it is out of contrarian, that the fear what they have worked for so long may not be useful after all. But true researchers should feel excited for the opportunity to be able to participate when the theory is still vanilla and contribute to it.
Of course, more research in tools for model interpretation would be awesome, and my own lab has done a lot towards it, and this remains an important topic. More is desired, but what we have right now is pretty good too, and is not at all inferior to old-school methods, esp. considering the performance.
Do you know of any work on interpreting neural nets that are being used for non-image tasks?
I think you are right on that one. The Standard Model for example is absoluteley not elegant, it is a huge convoluted mess of parameters and constants (just google for "Standard Model Lagrangian"). Yet, it is the best we came up with, explaining the dynamics of fields and particles[0], that surround us. Correct answers don't have to be elegant per se, its nice when they are, but it's not a prerequisite.
But, consider this: You can train a Neural Network to predict the distance of an object, thrown with velocity v_0, an angle of α, under the influcence of gravitational acceleration g. After a few hunderd rounds of training, a suited NN can reasonably predict the outcome of said experiment, with the input: v_0, α and g. And, as you pointed out, a researcher can explain to you why this is the case based on activation, feedback-loops, learning algorithm and other parameters. But neither the NN nor the researcher will be able to give you a rule like: "F=(G m_1 m_2)/(r^2)" to explain the underlying reasons for the object-trajectory dynamics.
A Neural Network can give you answers and predictions, yes, but you are not able to incorporate them into a wider theory, since the output is always numerical in nature. It is also always tied to one specific fact you are interested in and cannot give you a generalization.
[0]: In the model of QFT particles are also fields of a different kind.
That means one should expect them to perform tasks about as well as a mouse brain can do them, between a third and halfway to adulthood.
And yes, that level of cortical processing does not seem to support coming up with symbolic systems and attempting to use them to describe the world around them. Absolutely.
Mice don't do that either.
Come on, you can just fit the NN numerical outputs with Excel and come up with tentative explanations for both the underlying elegant rule and the science. The point with NN-based science being you are not left guessing much, instead retrofit the NN numerical output and build a coherent narrative with other accepted results.
On the other hand...
I'm not sure if there's been work done in the domain of pattern recognition, but here's an overview where in an example, Navier-Stokes is identified from data:
https://sinews.siam.org/Details-Page/data-driven-discovery-o...
Maybe a second NN could find that for you. You might end up with factors like "speed of an unladen swallow", but perhaps all that is needed is some nudge away from that local maxima to keep searching the space.
And sure, it doesn't output a nice simple equation. But in most real problems, there aren't nice simple equations to output. Physics is sort of a fluke, where the fundamental laws of nature happen to be so simple we can do that. In something like biology, you are never going to find a simple equation that explains a biological system. They are complicated messes of millions of interacting pieces. Even if we knew how everything worked, we don't have the computing resources to accurately simulate it anyway.
And we are actually seeing neural networks be useful in these areas. They are able to predict what chemicals are more likely to be useful drugs. If you restricted yourself to using only simple mathematical models, you would never get these results.
This very article is about the very difficult domain of image processing. Why should we expect the statistical distribution of photographs to be simple and explainable by simple math? Finding an equation that accurately models the shape of a human face is pointless and misguided. Before NNs came along, they were doing crazy stuff like handcoding complicated heuristics that detected dark spots that might be eye sockets. There's nothing elegant about that. Physicists are really blessed to have a domain where they can ignore stuff like air resistance. Where the data really does fit simple equations.
And as the parent comment says, you can interpret neural networks. You can observe the activation behaviors, and discover what features it's learned to detect eye sockets, or chemicals with certain properties. You can see how the output changes as the inputs change, and fit a simpler model to that.
If you are simply amassing data without purpose then what you could do is to try to visualize the layers of the network to see if any surprising features turn up but that will only work if the features are obvious enough to stand out and if that were the case I would suspect we would not have this discussion in the first place.
Computers excel at: remembering stuff forever and speed.
So any kind of improvement that a neural net or any other solution would bring to the table over a human would likely fall in either one of those categories, either the computer is faster at solving the problem, or its ability to remember and apply a vast amount of data to the problem will give it a (slight) edge over what a human could do, or maybe it will be just enough to reach parity.
It will not tell you what is and what is not significant in the input, though, with enough samples you with your slower but much superior intellect might be able to draw new and far-reaching conclusions once you are exposed to the data in sufficient quantity yourself that a pattern suggests itself.
That's in a way a very nice collaboration between man and machine, each doing what they are best at.
A key element in something like this being at play would be the computer consistently being at odds with experts in the field but being right more often than not about those cases. That would be an excellent opportunity to wake up to the possibility that there is something that should be obvious and noticeable but that still got missed.
> So any kind of improvement that a neural net or any other solution would bring to the table over a human would likely fall in either one of those categories, either the computer is faster at solving the problem, or its ability to remember and apply a vast amount of data to the problem will give it a (slight) edge over what a human could do, or maybe it will be just enough to reach parity.
> It will not tell you what is and what is not significant in the input, though, with enough samples you with your slower but much superior intellect might be able to draw new and far-reaching conclusions once you are exposed to the data in sufficient quantity yourself that a pattern suggests itself.
I think that its too early to say how AI will be deployed in practice -- whether it will augment or replace human roles. The ambition is certainly to produce enough "intelligence" to supersede a large fraction of human decision making.
As we seek to make computers more powerful, robust, and versatile ("AI"), it seems like we are pushed towards more organic computational structures. If the trend continues, it would imply that the closer we get to AI, the less their strengths and weaknesses would resemble computers of yesteryear. The interesting possibility is that one might be able to have an interpolation between the strengths of humans and computers.
But this is really an optimization problem: representing your formula as a bunch of weights that are then used to drive a cascade of multiply-add rules to derive some output is an imprecise but roughly accurate way to model a problem. That there is a much more direct and analytical way of modeling that problem requires intelligence of a different order.
But a program like mathematica gets awfully close to that, so that's not a problem that is a good fit for being solved with a neural network even if it would work.
I'd prefer to use a neural network for problems that have a less well defined solution.
It is not impossible. There is a combination of a special neural network architecture and strong sparsity-inducing regularization that makes it possible to learn equations from dynamics dataset: https://openreview.net/pdf?id=BkgRp0FYe
If a powerful idea is not elegant, we have the option of creating new mathematics in which it _can_ be elegantly expressed.
We may not quite be there but in terms of the reasoning facilities of our minds, we seem to have understood nearly all that can be understood about the universe. All the quantum stuff is simply not understandable using all that machinery but seems to require we run a virtual-machine like thing grafting reason onto weird probabilistic models. Given all this, I don't immediately assume the most "elegant" solution is really the best anymore for most of the big problems we're now trying to solve. At least in the domain of science.
One example is in the child services. We've got enough digital historical data from child services, that we're able to build models, that let's us predict whether or not a child will have to be removed from abusive parents 5-10 years from now.
This information is relatively useless though, because there is no way to explain how it works in a sense that the general public will accept.
The public is probably going to adjust eventually, but until that happens, you'll see an increasing push for documentation on how a result was reached.
Ironically an algorithm that explains how it ended up with the results that it did, isn't any more moral than an algorithm that didn't. I mean, you're losing you autonomy and freedom when the government knows what you'll do before you do either way.
Not sure what "moral" means, but two things:
1. If black boxes make unexplainable decisions, that gives too much power to the person who reads the box's answer,
2. If the child services algorithm explained its prediction involved e.g. alcoholism, then the parents could try to quit.
Mathematically, nn training is an unproven algorithm, it just works empirically surprisingly well. It's analogous to the simplex method which worked well for years without theoretical justification.
Older methods were proof heavy and required more mathematical sophistication but once you understood them, you could easily implement the idea. A lot of deep learning is comparatively mathematically simpler but there is often much incidental detail and folk knowledge that ends up being important but not mentioned. It is in this sense that one can call DL inelegant. A lot of that is because so many are racing to publish results, there's not much time to contrast with prior art or do much more than justify with often very expensive experiments with hastily scrawled descriptions. Lots of ideas generated without much context means many of them are quickly dropped and forgotten, sometimes without sufficient justification.
> Examples of elegance in deep learning
GANs, VAEs, gating, convolutions and weight sharing are indeed great ideas. IIRC Hinton's early papers did inspire Friston's influential work in theoretical neuroscience. However, translation invariance is double counting the advantages of convolutions. Policy gradients, the log-derivative and reparametrization tricks are independent of deep learning. Variational inference is more due to how far reaching that idea is if you want efficient generative models. Game theory...well lots of ideas are connected to it, including evolution and the dead simple weighted majority algorithm. Split brain is really reaching, especially since it's now looking to be one of those ideas that will need a decent amount of revision.
That said, I agree that too much emphasis is placed on interpretability. If a function is too complex, then a human will simply be unable to fit it in their working memory. Nonetheless, the ability to introspect on some of a neural net's decision is vital. As you say, visualizations and mappings which project to a simpler function space but preserve most of the detail are just as possible with neural nets as with any other method.
In this post I might come across as disparaging of deep learning, but this is certainly not my intention. There are lots of excellent papers which elegantly tackle difficult questions. They ask: what invariances and structures do we seek? How do we learn good representations we can sample from? How do we keep learning stable and achieve good gradient flow? How can we capture longer range correlations? These also offer insights that are not limited to deep learning. You just don't hear as much about them because they are not as shiny.
And if you're looking for mathematical elegance that cuts to the heart of the matter, specially with the rising importance of generative models, you could hardly do worse than start from anything written by Shun-Ichi Amari.
People always want to jump right to the shiny stuff and skip the basics, this is understandable but suboptimal in the long term, in any field. But a good compromise on this matter is the excellent free deep learning book by Goodfellow, Bengio and Courville.
But really. Secretly we all expect one to exist below it all.
[1] https://physics.stackexchange.com/questions/34217/why-do-peo...
[2] https://blogs.scientificamerican.com/critical-opalescence/do...
LSTMS amaze and dumbfound me in equal measures and I admire anyone who can understand them, at all.
Could you please indulge my cluelessness and explain to me what is mathematically elegant about the gating mechanism in LSTMs and how it adds to interpretability?
0: http://r2rt.com/written-memories-understanding-deriving-and-...
The OP made a big todo about how the article above is misguided and the person who wrote it has a poor understanding of deep learning, so I'm assuming he or she has a very good understanding of the subject.
I mean, I'd hate to believe the OP is just playing deep learning bingo for HN karma, you know?
I know, right? Like, take this model I trained this morning. Here's the parameters it learned:[0.230948, 0.00000000014134, 0.1039402934, 0.000023001323, 0.00000000000005]
I mean, what's "black-box" about that, really? You can instantly:
(a) See exactly what the model is a representation of.
(b) Figure out what data was used to train it.
(c) Understand the connection between the training data and the learned model.
It's not like the model has reduced a bunch of unfathomably complex numbers to another, equally unfathomable. You can tell exactly what it's doing- and, with some visualisation, it gets even better.Because then it's a curve. Everyone groks curves, right?Right, you guys?
/s obviously.
[Full disclosure: that comment was originally mine from a different thread; I'm not affiliated with the OP in any way.]
I probably should have cited it (was on mobile, but oh well :/ ). Well, I'll do it now: https://news.ycombinator.com/item?id=14219450
If anything, please consider this copy/edit as a sign of respect, and not of mockery :)
Many of us, children of the personal computer revolution, were attracted to computers for its empowering and democratizing effects. Just with a computer you could build anything! You are as powerful as any of "them"! The proprietary software model was developed to monetize end consumers by enchaining them, but the free software model and the fertile communities of the early Internet showed that the model of personal computation could not simply be displaced by legalese.
But if we take the author's position on the impact of deep networks, the model of computation is changing again.
If teaching deep networks is the new way to write useful programs, our brains and personal computers are not enough... they are obsolete. We now need clusters and, most importantly, massive access to data! This gives extraordinary leverage to the hoarders of data and servers of the Internet, the googles and the facebooks.
I'd like to think that maybe we are just at the beginning, the early "mainframe neural networks" era. That we just have to wait until there is enough technology and new markets discovered to build the "personal neural network". That consensual, open and distributed ways of sharing data will emerge and that new massively parallel computers will become affordable by the masses. That the models of neural networks will become well divulged and simplified and kids will be able to program them with their "Neural Basic"...
But at the moment the prospects don't look good. The Internet is becoming more and more centralized, personal computers harder and harder to program and there is a general "war on general computation" [1]. Even universities seem to bedisplaced by the Internet Lords in driving neural network development...
Maybe it is already a good time to start thinking about what "Libre Neural Networks" look like. And how can we get there.
It's important to remember that the intended product of science is not tools, it is understanding. Deep nets may produce very elegant and interpretable tools, from a very elegant and interpretable theory, but that is not what the scientists are looking for. They are generally creating and assessing theories in their subject domains, which are then evaluated by the tools they have built.
Deep learning does do a great job of providing a baseline ("your theory must be this accurate to matter"), but it seems to do a much less great job at extracting new understanding.
Having said that, there still exists the lunatic fringe (it's a compliment) inside deep learning who continue to work on the task of generative modeling in the hope of "understanding the world" rather than "solving a task". Yoshua Bengio and Yann LeCun don't miss an opportunity to impress on the world how important unsupervised learning is.
Physics too has abandoned understanding for predictive power. If you take philosophy of science seriously, you either have to pick "science is whatever the elite society believes in (Kuhn)" or "science gives us tools for prediction (Post-Kuhn)". Claiming that science helps in finding truth or understanding, is going to put you in a very indefensible slippery slope.
"Understanding" is a subjective human concept, and not worth pursuing. Humans evolved to run from tigers, and struggle with harder tasks like understanding quantum mechanics or understanding the brain. The directionality of scientific progress as evidenced by QM is not "understanding", but tooling and predictive power. I hope Deep Learning is going to follow the path of predictive power, rather than regress into the pseudoscientific mess of elegance and understanding.
Isn't human understanding of how our environment works precisely how you're able to sit in your chair and write comments on HN for others to read?
You might slap a few pieces of wood together and by brute force, create something like a chair and that might do, but it's very unlikely to be a good, safe, attractive, comfortable long lasting chair. Building quality furniture is complex and difficult work which takes a lot of skill, experience and understanding in order to do.
I understand your point, I never said there is anything wrong with experimenting with NNs, I just don't agree with the OPs sentiment.
This may be threadomancy, but I kept thinking about your comments and I realise now that there is a huge confusion between the different ways that the word "prediction" is used in the sciences.
What you mean when you say that you can use quantum theory to "make predictions" is that you can plug in some values to the formulae and come up with other values- the position or speed of particles and so on. This is a "prediction" in terms of stochastic determination of the state of the world within a probabilistic framework: you "predict" the values that some random variables will take.
On the other hand, what is more commonly meant by "prediction" in the sciences is the ability to anticipate novel observations, phenomena that have not been observed yet. For a famous example, phycisists and mathematicians predicted the existence of black holes before those were observed by astronomers.
This ability to foresee as-yet unseen phenomena is the true power of scientific theories, the context in which "prediction" is most often used in the sciences, and something that cannot be achieved without a thorough understanding of said theories. You can't take a black-box model of pictures of dog breeds and make a guess about what other kinds of dog breeds might exist, or might be created, than the ones in the original training set. The model may even be capable of doing that- but you, the person training it, are none the wiser. You can't use the model to expand your knowledge about the world. It's a closed loop.
So it's not really "predicting" anything in the way you say it, the way Kuhn meant it or anyone else means it.
(Edit: I knew I'd expressed this feeling before: https://twitter.com/copingbear/status/825098385548009472)
Drugs work but we don't know how a surprisingly large number of them work. Yet chemists don't complain about them. Theoreticians should not complain either. Deep nets give them a huge subject matter that probably hides many interesting insights in it. I mean, don't convolutional layers look like gabor filters?
The question of what can be understood from them can go deep, to the limits of the ability of language/math to express ideas.
If the task is interpolation between the data points this can be highly accurate. It the task is extrapolation such as prediction or designing a truly new machine or system in an engineering application, the approximation will often fail. The about one percent error rate in predictions of planetary motions from epicycles is one of the earliest cases of this problem.
The simple example is approximating data with a polynomial with an arbitrary number of terms such as in the Taylor series expansion. With enough terms a polynomial model can always approximate any data set arbitrarily well. Polynomial models can interpolate very well unless the data has some generally unusual behavior -- changing unpredictably at successively finer scales for example. However, finite polynomial models almost always extrapolate grossly incorrectly. As you move away from the data set in the space of independent variables such as X, the largest power N in X^N dominates and the polynomial approximation function blows up to either plus or minus infinity which is rarely physical.
What we think of as "understanding" corresponds conceptually in part to the ability to make accurate predictions. "Understanding" or "explanation" corresponds mathematically not to some arbitrary super-complex function with large numbers of arbitrary parameters but rather to a mathematical object such as as system of differential equations (e.g. Maxwell's Equations for Electromagnetism or the General Theory of Relativity) that express interrelationships among the variables and data points.
Firstly, mathematical elegance is not an objective metric by any means. Many mathematicians disagree on what results are most elegant, though there are many commonalities as well.
Secondly, all of machine learning is quite elegant. I'm by no means an expert since I only took two courses in my fourth year that were surveys of many methods (quite mathematically rigorous btw) but I've taken enough math to be able to judge what is elegant or not.
As the author himself says, he has slightly modified his research methods. These days you'd be a fool to ignore deep learning, whether you have a deep understanding of it or not.
"Elegance" can indeed not be defined, but I believe the misunderstanding here is more about the point of view: Deep neural networks are plenty elegant when looked at from an ML POV. But, if you're an expert in, for example, image processing, using a DNN as a tool, the solution it may give you won't be "elegant" by anybody's definition. It is, after all, a repetitive formula with X million arbitrary floats.
Previous ML methods usually resulted in models, or formulas, that were small enough to "grasp" intuitively. The "beauty", or "elegance" was that often, you could find connections from terms in your model to the real world.
Which is not true in practice. SVM in reality is an instance based method, with thousands to 10 thousands high-dimensional vectors as its 'parameters'. Random forests, which is famous for its ease of interpretability is no easy meal either when you end up with 100s of trees with outrageous branches. Not mention that real world model are embarrassingly complex ensemble of smaller but still complex models.
As a tool for basic research, particularly in the biomedical realm, it's awful. You get a system that performs pretty well most of the time, but tells you nothing of interest.
That's fine, as far as it goes, but "performs pretty well" is so much more useful in industry than "we know how X biological system works" that entire fields that are interested in the mechanics of how perceptual biological systems work and may be fixed are getting eaten alive. To our collective detriment, I think.
Yet this paper [1] from 2014 works extremely well, no neural networks. Here's one from 2008 [2] that worked extremely well too, better than most early NN techniques.
Although, deep learning is far from the intuitive approach that existed 10 years ago. It's clearer now how to reason about models, layers, activation functions etc. As much as there's no mathematical foundation, I do not believe it helps that much in the case of SVMs or whatever else. There still has to be experimentation, approximation and proper testing.
[1]: https://dspace.mit.edu/handle/1721.1/100018 [2]: http://link.springer.com/chapter/10.1007%2F978-3-540-88690-7...
These are actually excellent examples of how deep learning has totally changed our expectations of what can be done with computer vision by using deep learning.
The hype was nowhere near just because there's no deepness.
I admit that https://arxiv.org/abs/1705.01088 this beats everything and is extremely powerful and simple but the hype for deep things seems a bit too strong.
There are also a few methods, like boosting, that are purely the result of theory.
I'll take the other side of that bet.
I also think it's unscientific to expect mathematical elegance from everything. Math is ultimately a logical descriptor tool for AI, a means, not an end in itself. Besides, there's actually a host of involved mathematics underlying the seemingly simple SGD of deep nets. For example, tropical geometry has been used to analyze the loss surfaces of ReLU networks and random matrix theory has been employed to analyze the loss surfaces with respect to the quality of local minima.
Tl;dr: deep nets are far more interpretible than they've been given credit for and they're also more mathematical than some (such as this author ) would have you believe.
Computer Science is full of abstraction and complexity. Denying it to machine learning seems conservative and naive at best.
One easily-described approach that is state-of-the-art in a massive number of domains? That's pretty f'ing elegant.
1) Optimization theory - which says, if I repeat this iterative method N times I will find a (nearly) globally optimal solution
2) Statistical theory - which says, if I observe this process N times I can accurately estimate a population quantity with high probability
Deep learning does not benefit from the same theoretical guarantees.
For the most part the response from the community is "but it works really well!" which is a fair and valid response especially since what most practitioners care about is predictive accuracy.
Personally, I find applying neural networks extremely annoying at times due to the amount of twiddling of hyperparameters, slow convergence, etc.
I don't agree with this statement, a simple look at cs231n lecture series will show you how much math is involved. A lot of articles/people claim its a black box, but while writing a small network architecture you realise it's not. Methods like stride, padding, activations, learning rate, optimisations, drop outs etc. give you "aha" moment which is followed by a mathematical explanation. One should study the topic thoroughly before criticising it.
A hypothetical example: A trained decision tree algorithm for a medical decision may place the gender, or the age, of the patient at the root. That lends itself to a quick interpretation as to the relevance of factors for treatment, whereas with a Neural Net, you'll get millions of arbitrary floats that do not impart any meaning just by looking at them.
That's not to say NNs don't sometimes give tantalising insights, as the article points out. I've seen a view visualisations of generative character models that were fascinating–such as finding individual neutrons tracking sentence length or nesting depth. Same for some interesting patterns emerging in the intermediate layers of object recognition networks: oh, I never knew ears were so important for face recognition.
Also, what are we approximating? Continuous functions? Non-continuous functions? Are they even functions and not probability measures? Are those functions arbitrary or do they represent something like a manifold?
And the most important: how well are we approximating whatever we want to approximate? The universal approximation theorem gives uniform convergence for measurable functions, but do not specify at which rate or depending on which parameters. It is a strong theorem but not that surprising from the mathematical standpoint, where you already know that you can approximate any function by continuous, compactly supported functions.
Finally, how do you mathematically define the problems that arise in neural networks? What is overfitting? How does the learning algorithm affect the results?
The fact that some techniques are justified by mathematical explanations does not mean that it is mathematically elegant. For it to be mathematically elegant you should have at least clear definitions of the objects of study and the problems you want to solve. I don't think this is the case in neural networks.
I disagree. Elegance and interpretability comes with one big flaw: assumption, lots of them. While deep learning based methods assume few or not. If they are so elegant why would them underperform? Probably it is because the problem we are trying to solve doesn't follow the assumptions, like convexity, etc.
In that sense, one can say deep learning methods are even more elegant, because they can work end-2-end. In that sense, NN can be a blessing, because now we got the answer, we don't need to assume anything, but just to DECODE it.
Agree. Not just end-2-end, but the same model can be applied to many different tasks. DeepMind's recent papers make progress in image generation, audio generation, language modeling & machine translation using the same basic construct of a masked dilated shifted convolutions.
One tradeoff might be that with deep learning, it's unlikely to fail, but you won't know when it will. With conventional methods, you may not get as good of a performance in comparison, but have a better sense of the domain of applicability because you have a better understanding of the feature space (for instance, you can determine similarity of new data in this feature space when it arrives to determine whether it is similar to past data your model has been trained on).
No doubt that DNNs are subject to the same theorem, the only question is, how does their hypothesis space looks like? Does anyone have idea about this? I suspect we don't really know what the DNNs assumptions are.
Yeah, but that's tautological. What I mean can somebody sit down and for a given DNN architecture, write down (at least approximately) the set of functions that it can learn? Or more importantly, what functions it cannot learn? Or at least, how many bits are assumed and how many bits have to be learned?
I think that is what bothers people about DNNs. I personally think they are sometimes even inefficient - we are learning them much more parameters (bits) than the actual hypothesis space requires.
For a two-layer architecture with ReLU activations and n units in the hidden layer, this is the set of piecewise linear continuous functions with n kinks.
The architecture of NN itself is the biggest assumption here, as it is how it will be used to process the data. But they those operators are usually generic, a lot of them just umbrella operators for matrix multiplications + non-linearity in between.
I have a really hard time keeping reading after this passage.
Exam score= cat_a + cont_b + err
but how do you say something like cf male female has 10 score advantage and family income impact is ...
Can Galileo invent physics by doing nn
This is the telegraph operator lamenting all the time spent learning Morse Code when the telephone was invented.
I think scientists can be motivated by looking for new ways to solve real problems. Don't weep over the time spent understanding the classic models. Rejoice that we have a much better tool to further mankind.