Using neural nets to recognize handwritten digits
neuralnetworksanddeeplearning.com
neuralnetworksanddeeplearning.com
Couldn't "function approximator" describe most machine learning approaches? And a nice fitting algorithm is of course the goal.
I'd go further - the "nice fitting algorithm" is to minimise the error on the training set as a function of the weight parameters, and one obvious way to (locally) minimise that is gradient descent + the chain rule.
The math / applied math machinery is all incredibly general, and a useful way to think about many machine learning algorithms.
http://en.wikipedia.org/wiki/Gradient_descent http://en.wikipedia.org/wiki/Chain_rule#Higher_dimensions
Not to say that applied math is the only valuable perspective to think about, there's clearly also statistical and computational views as well. E.g. we're trying to approximate a function that we can never directly evaluate (error on the samples we havent seen yet)
Looking forward to reading chapter one when I have time, though I suspect it will confuse me quite a lot...
edit: I see he is, already commented before I wrote this
I think this comment from the article needs caveats. Of course, a neural network would not qualify as Turing Complete just because it's finite. Keep in mind also that neural network, lacking anything like counters, tape, or recursion, couldn't approximate a Turing in the way that a finite Von Neuman architecture machine does. (A NN can represent any given function over a domain if it get large enough, kind of the universality of a finite automaton).
I know this a reference to this generation of NN having overcome an earlier problem of not being able to represent a NAND gate but still, it's worthing keeping mind that an ordinary computer can simulate an NN with just a program but this doesn't work vice-versa, so that NN's in that sense are far from universal.
I agree that the relationship between circuit complexity and Turing machines is somewhat subtle, for the reasons you mention. The relationship is greatly clarified by the notion of uniform circuit complexity, which makes it possible to prove an equivalence between a (carefully defined notion of) circuit complexity and Turing machine complexity. Unfortunately, I don't know of a good online treatment of uniform circuit complexity. I learnt it through a 1993 paper by Andy Yao, but that's definitely not a good introductory reference!
In any case, in my book I'm using the term universal in the same way as people usually use it for circuits, i.e., it means the same thing as when people say that the NAND gate is universal for computation. Hope that clarifies things.
Can anyone figure out what each hidden node represents?
You can also select a node and press "A" (Gradient Ascent). This will change the input in a way that increases the selected node's value. By selecting an output node and mashing "A", you can run the NN in reverse, causing it to "hallucinate" a digit.
Of course, there's a tremendous amount going on, so my broader philosophy is to focus on fundamentals. Readers who thoroughly master the core ideas shouldn't have much trouble later getting up to speed with the result-of-the-month.
kernels learned by the first convolutional layer (the figure 3. on page 6) have uncanny resemblance to Gabor function-modeled orientation-selective cells ("bars and grating cell") in the primary visual cortex. Looks like computers are on the right track :)
http://www.cs.rug.nl/~petkov/publications/bc1997.pdf
"The discovery of orientation-selective cells in the primary visual cortex of monkeys almost 40 years ago and the fact that most of the neurons in this part of the brain are of this type ..."
The difference here is a "number game" - visual cortex contains cells whose receptive fields' positions, eccentricities, sizes, orientation, number of excitatory and inhibitory zones (e.g. Fig.1 in the link) make a reasonable coverage for the space of possible values. Ie. the number of these cells is in the millions vs. 96. Of course it is only a matter of computing power to run all reasonable combinations of kernels emulating the real visual cortex, yet it would put immense computational challenge onto the second and next layers until we understand what [should] happens there.
I'm not sure if I fully believe this, but certainly there doesn't seem to be a very principled way to choose your network architecture. Different people propose different ones, and the fundamental justification for each one seems to be: "look, we recreate gabor filters in layer 1 and we get good numbers at the end!"
Of course, NN people argue that that's almost exactly what vision people do as well, except in "feature-land" rather than "architecture-land".
well, i can see the temptation - the orientation and spatial frequency selectivity are the major characteristics of cells in V1 and the receptive field for the first layer there does look like Gabor
http://www.scholarpedia.org/article/Area_V1#Receptive_fields
I agree that such a good resemblance of the learned kernels to Gabor is too good, this is why i used "uncanny" :) If it is real then i think it manifests very interesting and, no pun intended, deep emerging properties of the neural net learning process (something along the lines "maximum entropy kernels while still doing the job" as the asymptotic state)
Btw, is it really selection or confirmation bias?
And to expand on previous point of convoluting the input with many-many kernels - happens to be at the order of 40 per "pixel":
"V1 contains a vast number of neurons. In humans, it contains about 140 million neurons per hemisphere (Wandell, 1995), i.e. about 40 V1 neurons per LGN neuron. Such divergence gives scope for extensive processing of the images received from LGN."
I was able to get everything by looking you up on github and using the url of the repository.
Edit: Also, you might mention the repository earlier, because it's rather large and I've had to break from the book while it downloads.
https://github.com/bcuccioli/neural-ocr
There's a paper in there that explains the design of the system and my results, which weren't great, probably due to the small size of training data.
The training data is the MNIST dataset released by NIST some time ago, anyone is allowed to use it. It truly is no surprise to see it here, as it is very a very commonly used dataset in ML tutorials/books. It receives some discussion in Artificial Intelligence: A Modern Approach by Russel and Norvig, and even in the Theano getting started tutorials.
In fact, if the goal of the book is to educate people on the field then I would say it's definitely best to use the standard benchmarks. It lets readers relate what's in the book to the literature, should they desire, and just like in academia, it lends credence to the author's statements. I've seen people take a lot of flak for publishing writings that don't use the standard datasets, but make claims of progress.