A Gentle Introduction to Backpropagation
numericinsight.blogspot.com
numericinsight.blogspot.com
[1] http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/
To summarize, people generally abandoned backpropigation trained neural networks for Support Vector Machines because neural nets require labeled and limited datasets, and work slowly and especially so when dealing with multiple layers which is sort of the whole point.
In my work in JavaScript, I was able to pull off only a single layer perceptron and it is neat but limited in what it can model.
In fact, deep neural networks are trained in an unsupervised manner at first, but then back propagation is used to "fine tune" and improve the results. Because they require unlabeled data sets, and can perform so well, research into neural networks has experienced a recent resurgence.
By the way, any talk by Geoff Hinton is fantastic. If you are interested in neural networks and their capabilities, and you haven't already seen it, his Coursera course [1] builds from a simple linear perceptron to the current deep learning methods.
[1] https://class.coursera.org/neuralnets-2012-001 (You'll have to sign in to see it)
Neural nets are from a silver bullet, and shouldn't be used where feature introspection is a huge requirement (this is why decision tree/random forest is popular), but they are far from being what they were in the 90s.
Note that I have a commercial interest in this so there's going to be inherent bias in my opinions.
To be fair, javascript isn't a scientific computing language. To do most neat training with neural nets, you're going to want to either scale them out, add more layers, and/or use GPUs. That being said, a neat toy example in javascript is convnetjs[1].
He uses a nice graphical approach that is easily understandable yet formal. It's been many years since I've read it to learn for an exam at university but I remember it was an enjoyable read and I wished I had more time to spend on the book.
[1] http://www.amazon.com/Neural-Networks-A-Systematic-Introduct...
It has a similarly approachable yet formal style, and I have recommended it to people with no ML experience in the past who have found it very intuitive.
In the context of CS what's the difference between "learning" and optimizing?
But I would say that learning is a higher level concept. So you would use an optimization algorithm to learn the solution for an SVM from data.
Learning in this case generally means the creation of models. Optimization is the process by which this happens.
A few examples:
Unsupervised learning - typically clustering and grouping things
Supervised Learning - Example based learning where you're trying to label something. This could be sentiment, spam classification, object recognition over pixels, even extending to sequence labeling.
Prediction/Regression - Predicting values based on a learned function.
Learning is more about the goal trying to be achieved. Optimization is more of numerically solving relative to an objective function (also called an error function)