First, the term "deep learning" does not only apply to artificial neural networks, although that's what it's usually used as short hand for. It can also be used for reinforcement learning, for example.
Second, "deep" is a technical term. In a neural network, it refers to the number of hidden layers. With NNs, you have so-called visible layers on the edge of the net for the input and the output. In between them you have the "hidden layers" stacked up. Deep means more than one hidden layer. So a shallow neural net of one hidden layer is not deep learning.
Each of those layers is composed of nodes. And as the data passes through the net, it is recombined at each layer in ever more complex ensembles which are given different weights, or importance, with regard to the final decision the net makes (classification, clustering, regression).
Deep NNs do two things that make them extremely powerful in relation to other machine-learning algorithms. They extract features from the data automatically and assign them importance in the process of learning. They do that using backpropagation over many hidden layers.
Here's an example: Let's say you have an image and you're going to classify it using a feedforward network (that's the simplest kind of network, because the data simply passes through from one side to the other once and is scored).
It's a digital image, so its pixels are just numbers, and those numbers can be flattened and arranged in a single column, so you just have a vector. That vector is passed to the net. Each pixel's number goes into a node of the net's input layer. (If your image was 28px*28px then you'd have an input layer of 784 nodes.)
Each of those pixels is a feature of the image. But when humans recognize images in photos, they rarely examine the a photo pixel by pixel. And many photos might have the same sorts of pixels -- red or white or black -- but they would differ because the pixels would be arranged in different patterns.
Those patterns are more complex features composed of simple pixels: a line, block or curve... In a neural network, those complex patterns are represented by the hidden layers. Each hidden layer represents increasingly complex features, based on the combinations of the layer before. So we would graduate from pixels in the input layer, to edges in the first hidden layer, to intersections of edges in the second hidden layer, and so forth until maybe we have nodes deep in the net that represent a nose or an ear extracted from a photo.
This is called feature hierarchy, and it's one of the central characteristics of deep nets' structure. You can't do that with shallow nets. And deep nets allow simple features to recombine into more complex ones automatically. That is, you don't have to waste time with the feature engineering of most machine learning algorithms. (You waste time tuning your net's hyperparameters instead ;).
A network learns how to assign the correct importance to various features via a process called backpropagation. Backpropagation is like reflux. You take something, and you send it back the way it came. In the case of neural nets, we're taking error. That is, every neural net starts out stupid, and it succeeds in learning something, it becomes less so. The measure of that stupidity is its error. And you calculate the error of a net by taking the guesses it makes about the data it sees, and comparing those guesses to a ground-truth label. Let's say you have a ton of images properly labeled as ship, cat, flamethrower, tiger and schoolbus.
Each of those labels is one output node on the net. The output nodes represent the final combination of features, which coalesce in a label we care about.
You pass the images into the network -- which operates behind a veil of ignorance and doesn't know the labels -- and you ask it to fire the output node that it thinks is the right guess. It does that, and almost always gets it wrong. Let's say it sees a photo of a cat but chooses flamethrower. You take the difference between those, some measure of error, and you backpropagate through the nodes of the network.
Each node in the hidden layers uses coefficients to assign importance to input. That is, it's multiplying the numbers that pass through, in order to amplify or dampen them. Those coefficients, or weights, determine what guess the net makes about the data in the end. They have a causal relationship with the net's error. So if you change the weights, you can change the net's final decision. Backpropagation involves going to each weight, establishing its relationship with the final error (using a partial derivative), and then modifying the weight slightly in order to decrease the error.
Once the weights are adjusted, more data is passed through the net; it makes more guesses; the error of those guesses is measured; and that error is used to adjust the weights. This goes on for a long time, basically until you can't shrink the error any more.