On a superficial level it seems like it:
1. Generalizes deep learning to an optimization function on decomposable input, and
2. Reduces the number of parameters required to learn the input by exploiting the structure of the input, thereby making learning more efficient.
Is that correct? Is it completely off? What am I missing? Is there any more meat to the article than this?
Could someone who has upvoted this (and ideally understands the topic well) provide a different explanation of the concept? It would be great if I could see a real world example (even a relatively trivial one) represented in both the traditional matrix computation form and the sexy new differentiable form.