Standard feed forward deep net: like the ones you used at university but with a few important features. One, you can stack layers on top of each other then train the whole network with back propagation. The nonlinearity you use can be important (rectified nonlinear units--Max(0,x) are popular now). Regularization can be important (with dropout being a popular method). Pre-processing your data can also be important (eg scaling the inputs, subtracting the mean, ...). How you initialize your weights is important. Using tricks like momentum and learning rate decay are important too.
Autoencoder: A standard feed forward deep net that tries to output an "uncorrupted" sample of a "corrupted" input. So basically, take an image, add noise, ask it to output the denoised image. Why you ask? Well, in doing this, the hidden units (the layers in between the input and output) tend to discover new features which can then be used elsewhere in machine learning pipelines. The big win here is automated feature engineering.
RBM: Restricted boltzman machine, sort of like an autoencoder but not. I'm not convinced one is better than the other, but there are definitely differences.
Recurrent neural networks: Like standard feed forward deep nets but extended to time series problems. The hidden units at each timestep feed in as additional input to the next timestep (along with the new input). So basically, each step it gets this as input: [whatever came before described as hidden units] + [input at T]. These are my personal favorite and currently hold state of the art in speech recognition, at least academically.
Convolutional neural networks: State of the art in vision, this one is kind of hard to explain quickly, but you can think of it kind of like this: I take some image where I want to recognize objects in it. Then I train a bunch of neural nets that represent different aspects of an image (one net might look for vertical lines, another might look for horizontal. It figures out what it's looking for automatically and distributes what it thinks are important across these little networks. You can think of this as "feature detection") You then take a window on the image (400x400 image, maybe you've got a 20x20 window) which you then slide across the image and check each section for the presence of whatever the net is looking for (vertical/horizontal lines, etc.) There's several layers of this (specifically a convolution operation followed by a pooling operation) before the result is fed to a standard fully connected feed forward net which then outputs a prediction.
While lower layers look for low level aspects of an image, each layer progressively looks for higher and higher level phenomena--for instance, when you hear about "the cat neuron" it comes from probing a neuron in the higher levels of a convolutional net and finding that pictures of cats happen to turn this neuron "on" while pictures of anything else don't. What is "high level" exactly? In this case, it really just means composed from "lower level" components... Some also say "higher level = more abstract" but I don't know this is quite the case--abstract means something different to me.
An important consideration is that the conv net architecture allows you to cut down on the total number of parameters (good when you've got something high dimensional like image data) and it also takes advantage of the fact that you can make certain assumptions about images and how objects move/appear in space basically.
Conv nets are probably the most difficult architecture to grasp and my explanation is extremely high level.
Ha, that was a bit longer than I had hoped, but I hope it's helpful to someone...
I might be totally off-base, but is Numenta's HTM (Hierarchical Temporal Memory) model a conv net? If so, there are some really good introductory/TED-talk-like explanations out there for the idea.
Personally, I think there are two big issues with what numenta was doing: they diverged too far from mainline neural net research (which, when they started in ~2005 was just before things started getting interesting--it was also a time when "mainline neural net research" was widely assumed to be at a dead end) and they tried too hard to come up with something that was biologically plausible rather than mathematically expedient. Sort of how airplanes need wings but don't need feathers. As far as I am aware, HTMs really just don't work well in practice (and by that, I mean they are not anywhere near competitive with any of the architectures I listed above).
You could, in other words, see a recurrent conv net as an opmitization of an HTM given a von Neumann architecture, or the reverse -- an HTM as an optimization of a recurrent conv net given a biological substrate (where it's much less costly to build tons of crummy processors and link them into an arbitrary graph, than it is to build a single fast processor.)
Again, though, I'm not an ML person, so I might be way off.
Numpy based neural networks are a lot more digestable as far as understanding all of the moving parts due to the easy syntax.
Neural nets are usually verbose if you use an object oriented representation. Things like feedforward and backpropagation don't require loops if you know how to think about individual neurons and weight matrices.
Also, for working with them, here's something that cleared up a lot for me: when initializing neuron weights, it's randomly initialized. It seemed like magic to me at first and kept tricking me up when getting started with them.
For those who do the JVM, I'm working on finishing up a library for myself to use in various bits of work I do.
Java based, but you get access to fortran matrix routines, and it's interopable with scala and the like as well.
https://github.com/agibsonccc/sda-jblas
For those of you who need hadoop or want to scale out neural networks, I'm going to suggest a friend of mine's iterative reduce project for neural nets on hadoop.
https://github.com/jpatanooga/Metronome
I just finished up a parallel training setup I'm going to release based on this for neural nets as well. When working with them, you'll discover crazy long training times. Multicore tends to make things easy if you know how to balance out the weights though.
I'd be happy to answer any other questions as well.
- Use the tanh activation function.
- Use many hidden nodes, e.g. 100 or 1000 for complex tasks.
- Tune the learning rate.
- Add a penalty of k*w^2 to each hidden node. Tune the parameter k to minimize the held-out error.
- If this isn't working, plot the error on the training set as you train. It should decrease to 0.
Read Yann LeCun's efficient backprop work for more tips.