Renormalization: A Common Logic to Seeing Cats and Cosmos
quantamagazine.org
quantamagazine.org
The write-up is a bit misleading, though: the model in their preprint[1] is about stacks of restricted Boltzmann machines (RBMs), and are very different from the other examples of deep learning mentioned. The Google Cat Detector model, for example, didn't describe a probability distribution over images, which is the kind of task that the preprint is about. And in almost all of the recent cases where deep neural networks have made substantial progress over the previous state of the art, the models have not been probabilistic, or trained layer-by-layer, or unsupervised, like the RBM-based approach in the preprint.
I don't speak physics, but my reading of the paper says that could be summarized pretty accurately as follows:
1. Hinton et al. (2006) showed that each layer in a stack of RBMs improves a variational lower bound on p(x).
2. Variational methods for RG also iteratively improve a variational lower bound on p(x).
3. The two methods would thus be equivalent, if we could fit them without error (which we can't).
4. Here’s a figure from a stack of RBMs that vaguely looks like RG results (not shown)
I don't see any comparison between the approximations that physicists normally use versus the contrastive divergence for training RBM-based networks, or evidence that their results are more similar in practice than any other technique.
Am I missing something?
[1] http://arxiv.org/abs/1410.3831 [2] https://www.cs.toronto.edu/~hinton/absps/fastnc.pdf
I didn't think of it to apply to deep networks, but it makes a lot of sense.
What is not so likely in the brain is that these topographically mappings are the only ones. There are probably some mix-and-match schemes that allow integration of information on longer distances on scales different than that of the most coarse layer. Just my two cents.
Oh yeah, and although physicists like fractal-like structures (the fixed point in the renormalization flow can be seen as having the same mapping between every two subsequent scales of granularity), people with real-life machine learning problems probably would like to study more temporal and dynamic aspects. When time gets introduced, things become complicated.
I haven't looked into the math myself, but apparently there are situations where the increase in representational power for networks of a given size can be increased exponentially by moving nodes into higher layers.
Even GPUs do sequential operations. You could never program the vast majority of algorithms like "if the input is exactly 100011110101... then output 100011001...".
If that seems silly, that is exactly how the proof that single layer NNs are perfectly general works. It proves that you can represent any series of if then statements like that in an NN. And people who don't understand the proof are mislead into thinking single layer NNs are just as good as deep NNs.