Deep learning
neuralnetworksanddeeplearning.com
neuralnetworksanddeeplearning.com
> A word on procedure: In this section, we've smoothly moved from single hidden-layer shallow networks to many-layer convolutional networks. It's all seemed so easy! We make a change and, for the most part, we get an improvement. If you start experimenting, I can guarantee things won't always be so smooth. The reason is that I've presented a cleaned-up narrative, omitting many experiments - including many failed experiments. This cleaned-up narrative will hopefully help you get clear on the basic ideas. But it also runs the risk of conveying an incomplete impression. Getting a good, working network can involve a lot of trial and error, and occasional frustration. In practice, you should expect to engage in quite a bit of experimentation.
There is a lot of "magical thinking" amongst people not actively doing research in the area (and maybe a bit within that community too), and I think it at least partly stems from mainly seeing very successful nets, and never seeing the many failed ideas before those network structures and hyperparameters were hit upon - a sampling bias type thing, where you only read about the things that work.
That said, any sort of hyperparameter optimization is extremely computationally intensive so random search is far from a panacea.
There are also somewhat better-than-random strategies such as Bayesian optimization and particle swarm optimization that can help you to search more efficiently.
16 days ago - https://news.ycombinator.com/item?id=9863832
8 months ago - https://news.ycombinator.com/item?id=8719371
a year ago - https://news.ycombinator.com/item?id=8258652
a year ago - https://news.ycombinator.com/item?id=8120670
a year ago - https://news.ycombinator.com/item?id=7920183
a year ago - https://news.ycombinator.com/item?id=7588158
two years ago - https://news.ycombinator.com/item?id=6794308
I believe that we are several decades (at least) from using deep learning to develop general AI.
I conclude that, even rather optimistically, it's going to take many, many deep ideas to build an AI.
The appendix linked there doesn't seem to be ready yet though. In any case, I like how this is phrased. I'd like to see some of the hype around deep learning calm down.
Out of curiosity, do many implementations of convolutional neural networks take advantage of FFT, DCT, or some other fast orthonormal transform to compute the transition between layers, or are the kernel sizes small enough that there isn't a great advantage to that?
They have a patent on it but did open sourced the code. They claim it's up to 24x faster than standard. But that is only true for an extreme use case, its only 2x faster on average.
It breaks even at 5x5 or so and gets dramatically better shortly thereafter. However, most of the convolutional nets in use rely on 3x3 convolutions because I guess reasons:
http://arxiv.org/pdf/1409.1556.pdf (all 3x3)
http://www.cs.unc.edu/~wliu/papers/GoogLeNet.pdf (3x3 and 5x5)
There's probably a new Imagenet winner in this somewhere IMO...
For example, the usual way to have a DNN learn rotation/scaling/translation is to do data augmentation and simply learn with all the data rotated/translated/shifted.
But there must be a way to have these input space symmetries reflect somehow in the structure of the network?
I tried googling this a bit but wasn't really successful - does anyone know whether this has been done?
My gut feeling is that the first convolutional layer's kernels, for example, would probably have a 'some are orthogonal' constraint due to this symmetry.
There have been papers about scale/rotation invariant convnets (again at the structure level) and also Networks that learn invariances without encoding them into the structure.
The former I am very interested in! Do you have any links?
Rotation-invariance is probably not really a thing you want. The visual world is not, in fact, rotation-invariant, and the "up" direction on Earth-bound, naturally-occurring images has different statistics than the "down" direction, and you'd like to exploit these. Animal visual systems are not rotation-invariant either; an entertainingly powerful demo of this is "the Thatcher Effect" (https://en.wikipedia.org/wiki/Thatcher_effect).
Reflection across a vertical axis, on the other hand, often is exploitable, at least in image recognition contexts (as opposed to, say, handwriting recognition). If you look at the features image recognition convnets are learning they are often symmetric around some axis or other, or sometimes come in "pairs" of left-hand/right-hand twins. As far as I know nobody has tried to exploit this architecturally in any way other than just data augmentation, but it's a big world out there and people have been trying this stuff for a long time.
I know that some translation invariance comes from e.g. the usual conv+maxpool layer structure, but there must still be several representations existing in the first hidden layer of the network stack, for the different translation shifts?
Especially rotation looks like something that should produce a lot of symmetry and shared parameters, but it also looks difficult enough for me that I rather would like to know about someone with mad math/group theory(?) skills who looked at that.
But thank you for the detailed reply anyways!
my reading for today, thanks for sharing!