Neural Network Architectures
culurciello.github.io
culurciello.github.io
If I could only read one thing to gain the technical grounding for this history, what should it be?
`Hacker's guide to Neural Networks` http://karpathy.github.io/neuralnets/
Read the 2nd one then.
It starts with with linear classification, then moves to neural nets, and then explains convolutional neural nets.
It introduces you to some of the underlying principles which haven't changed much over time. I highly recommend it if you want to get deeper intuitions on the principles of CNN, LSTM/RNN, Restricted Boltzmann Machines etc. Also, Hinton's Coursera lectures, though not sure if you can access it anymore.
http://deeplearning4j.org/neuralnet-overview.html
Also, this book is coming:
https://www.amazon.com/Deep-Learning-Practitioners-Adam-Gibs...
As mentioned in the article, using convolutional layers in ANNs was an idea from the 1980s, but networks that could be trained on the hardware available at the time were never all that competitive until recently. Once we figured out how to train big/deep networks (use GPUs, have lots of data, maybe use pre-training), CNNs started to perform really well. This did make a positive feedback loop: as CNNs started to work better, deeper networks in general started to get more attention, which got more people into CNNs, etc.
Sometimes you see this combined with a CNN. There has been a few question answering systems that have one or more CNN layers. In don't entirely understand these designs, but presumably the convultional layers are an attempt to understand the different orders of words.
There are lots of techniques that people use to try to make deep networks work well. Mostly theses are about making errors backprog better. One of the most successful recent innovations is the ResNet architectures (https://arxiv.org/abs/1512.03385), and the related highway networks.
Also convolutions are not only used in computer vision. For example, alphago used them: (the paper is called "Mastering the Game of Go with Deep Neural Networks and Tree Search"). In my opinion, I would say that convolutions should be useful whenever your data has a spatial aspect to it.
http://static.googleusercontent.com/media/research.google.co...
Processing power yes, but you can get started with a gaming pc.
But the post you replied to specifically said "as a hobbyist", so it doesn't really sound like there's much hope.
For example I was at a presentation where a person built a pretty interesting neural model based on 190k clinical records released via Kaggle. In most fields it is surprising how much data is easily accessible.
Why would it be against the original principles of LeNet?
So, if you're using 1x1 convolutions, I think you're basically having a neuron per pixel, so you're forcing your fully-connected layers to learn the spacial correlations of pixels, instead of capturing that information in a convolutional layer. In other words, you're wasting training on capturing spacial correlations of adjacent pixels instead of other correlations.
Saying "a neuron per pixel" doesn't mean anything, really, that way of thinking isn't helpful unless you're looking at small multi-layer perceptrons. The right way to think about things is that you have tensors and layers that compute new tensors from old tensors.
A 1x1 convolution only 'sees' the feature channels of a pixel, and does the same thing to each pixel. So a 1x1 convolution on a grayscale input (e.g. a 1x28x28 tensor in the case of MNIST) does nothing, basically, other than scale and bias every pixel by the same linear function. It doesn't "force the network to learn" anything, it's just totally pointless.
One of the uses of 1x1 convolutions is to collapse the feature dimension when you're deeper in the network (e.g. 100 channels to 10 channels) to reduce number of parameters subsequent layers need operate on. It's a "channelwise fully connected layer".
I think you're thinking of (and perhaps what the author was thinking of) is the practice prior to convnets of collapsing the image into a vector and then doing a fully connected layer on it. That indeed doesn't exploit translation invariance of natural images, requires the net to learn the same features in every required spatial position at great expense, and so on. But that has nothing to do with 1x1 convolutions.
[Also, sorry for attempting to answer your quesiton incorrectly. I was thinking of putting a disclaimer saying I hadn't worked with CNNs and so might be misunderstanding what the convolutions are doing; probably should have haha]
Maybe when the author was saying 'one can think the 1x1 convolutions are against the original principles of LeNet', he was anticipating my kind of confusion? :)
Correct. As I understand it, this would be applying a 1x1 covolution with 10 filters to a 100x512x512 tensor.