10B Parameter Neural Networks in Your Basement [pdf]
on-demand.gputechconf.com
on-demand.gputechconf.com
(My suspicions is that the initial layers can be "frozen" or trained separately, low-level features are pretty much the same for most images, also, maybe someone will figure out a way of merging several NNs in one, so you can paralelnize the whole training)
This presentation is only one small example.
Sure, there's a huge amount of research happening in the field right now. But you make it sound like there's a ton of low hanging fruit, which is emphatically not true.
Also, freezing the representations in the bottom layers usually doesn't lead to very good results. The representation the bottom layers learn in the standard formulation of deep feedforward networks is informed by gradient information in the output layers[1]. If you stop training the bottom layers after some time you're sacrificing representational power and in fact increasing the ultimate amount of computation that will need to performed to train the network to some level.
In fact, in a recent paper [2] some folks at Google describe how they achieved great performance training a huge CNN by inserting temporary output layers in the middle of the network while training (among other things). This increased the amount of gradient information that was propagated back to the bottom layers, forcing them to learn more powerful representations.
[1] https://en.wikipedia.org/wiki/Backpropagation [2] http://arxiv.org/abs/1409.4842
Yes, there are other methods. Contrastive divergence seems to be king right now - of note is Minimum probability flow learning [1] (of which CD is a special case of). However the flavor of these methods tends to be tuning the weights of the model in such a way to maximize how close the model comes to sharing the probability distribution of the data. One can generally not constraint the model parameters (ie by freezing a layer) and retain the models ability to 'learn' the data distribution.
I mean current GPU architecture is very good at vectorization, but is it possible to have a chip that has bigger cores than GPU has, but smaller that the ones on a CPU ?
It really seems that we have big cores because managing an OS requires it, but I wonder how realistic it is to make a computer that use more massive parallelism, there are many algorithms out there that can be alternative to sequential ones.