DeepMath Conference 2020 – Conference on the Mathematical Theory of DNN's
deepmath-conference.com
deepmath-conference.com
Heh. A burn both pointed and subtle. A+.
There is theory w.r.t to thick networks as well (e.g the link to Gaussian processes require infinite width).
Deep makes sense here.
Note that this is hortogonal to sparsedness vs density
In the theory literature, if you have a K-deep network, K=1 is the shallow case, K>1 is deep. Agreed naming could be better, but it's not like "deep work" or "deep thoughts" as the parent was stating.
Cancel culture will ensure that the conference doesnt happen in the name of fat shaming.
Deep Learning with Big Data
That stuff sells itself.
A sufficiently (infinitely?) wide shallow fcnn can work in theory, but is basically impossible to train[0].
It might be feasible to transform or 'compile' a trained deep network into a shallow & wide one, but I'm not sure there would be any benefits, absent sufficiently wide parallel hardware[1]
[0] Then again, deep networks were impossible to train for quite a while too, due to the exploding gradient problem.
[1] Although, the Cerebras Wafer-Scale chip does exist. Hmm.
Nnets have always been multi layer since they were invented. That’s the whole idea of progressive feature extraction, and the analogy with biological brain. Theoreticians referred to them properly as nnets or multilayer nnets. Later experimentalists simulated them, thanks to the availability of the computing resources, and experimentally verified that a multi layer nnet can be more efficient than a single layer one. They added superficial terms “deep” and “AI,” “singularity,” etc., which the media and tech industry amplified for obvious reasons.
https://en.m.wikipedia.org/wiki/Universal_approximation_theo...
I can give you a function that a shallow nnet would approximate better and functions that deep nets approximate exponentially better even with one more layer (in terms of number of neurons n). In the limit n->\infty, both reach arbitrary small errors (obviously often with different number of parameters).
For many years a 3 layer network (1 input, 1 hidden, 1 output) was the standard and considered sufficient, thanks to several relevant but non constructive approximation theorems (mostly one due to Kolmogorov and one due to Cybenko, if I am not mistaken).
They were also considered practically required because everyone was using sigmoids which have a vanishing gradient problem.
Several things were needed to break away to where we are now: unbounded functions (like ReLU) to avoid vanishing moments; a lot more layers; a lot more parameters and compute power.
When Schmidhuber (and later Hinton) showed that many-layer nets work well, that was non trivial and almost revolutionary.
It is now trivial and all nets are “deep”. But that wasn’t the case when the breakthroughs were made, and the terminology stuck.
Cybenko,Hornik etc studied the mathematical properties of multi layer feed forward nnets late 80s.
Obviously, in the field of systems control equivalent “deep learning” and “reinforcement learning “ were studied since 50s. This includes what’s called “back propagation algorithm”
It was all multi layer nnets until it was simulated.
I have no time to go look at all those sources now, but having dabbled in nets since the late '80s myself, I remember vanishing gradients were sort-of a surprise, because everyone was under the impression that simple backpropagation should just work, and it didn't.
A lot of that early work you refer to was also mostly about linear transfer functions, and though the exact type of non-linearity doesn't matter, some of its properties do - and as I mentioned, sigmoids - which were all the rage in the '80s - are a dead end with the wrong kind of nonlinearity.
Nothing about the structure* of multilayer models is new. But successfully training them - which didn't happen until Schmidhuber and Hinton (depends on who you ask ...) - is relatively new; and that advance is responsible for the term "deep learning".
We do not disagree about the details; but we do seem to disagree about the historical context and narrative.
I was unaware, but apparently Gribel gave a constructive proof in 2009 (link from Wikipedia article about KA rep theorem). I would have to read it and hope I am not too rusty to understand it before I could really ponder your question...
But I could offer two places I would have looked:
1. The approximation is of a continuous function, and such approximations (e.g. chebychev, bernstein) usually require that you be able to sample the function at specific points - but learning usually gives you training data that does not correspond to those specific points. It's possible that construction fails here somehow.
2. The approximation is too hard in practice. This is the too often the case for Breiman's beautiful ACE (Alternating Conditional Expectation) which, if you squint hard enough, looks like a two-layer network where each neuron has its own transfer function. The algorithm is incredibly simple in theory, but very hard to use in practice.