Predictive Learning [pdf]
drive.google.com
drive.google.com
Deep conv nets were not designed with prediction over time in mind. Here's one reason: Deep convolutional nets do not handle dynamical information from their lowest layers. By design, conv layers and pooling layers immediately begin discarding spatial information that could be used for building up predictions.
In contrast it is possible to start with recurrent/feedback networks at the very first layers of the network. These initial layers can begin building up predictions at the pixel, color, and lighting level (example: our recent preprint[1]).
My colleague, who as some of you know enjoys blogging, wrote a more thorough post in response to LeCun's recent CMU lecture on the same topic as these slides[2].
[1] https://arxiv.org/abs/1607.06854
[2] Blog: "A few comments on the Yann LeCun lecture at CMU, 11.2016" http://blog.piekniewski.info/2016/11/21/yann-lecun-cmu-11-20...
I just want to add a very simple statement:
In order to create a model of the world, the machine learning substrate has to have the capacity to cover the observed dynamics.
Hard not to agree with this, almost sounds like a tautology. Now let's try to derive conclusions:
- World dynamics is full of multi-scale interactions (e.g. in vision illumination of a single pixel depends on the whole scene and the whole scene depends on many tiny details). To capture that, the machine learning substrate has to allow low level representations access high level stuff. Hence feedback all over the place. Which is exactly what is seen in the biological cortex. This is not recurrent layer made of LSTMs, this is a FULLY RECURRENT system.
This will not be achieved with any franken-neocognitron deep network neither with MSE, nor adversarial nor even triple adversarial loss function. This requires a new approach and together with several colleagues after a few years of continuous and intense thinking and modelling we have proposed a solution:
http://blog.piekniewski.info/2016/11/04/predictive-vision-in...
As well as a full paper https://arxiv.org/abs/1607.06854
Now, I'm not saying this is all done. This is just a beginning of a really exciting research adventure and many things look very promising. It will require however for the AI field to get out of a pretty "deep" local minimum it is in right now.
The interesting question, then, is how we make machine-learning systems do prediction well. Probabilistic/generative models have been "wandering in the desert" for a while now because Monte Carlo methods are just so slow, especially for high-dimensional, hierarchical prediction problems like we want to solve in machine learning. On the upside, STAN now has automatic variational inference for continuous probability models, and work on things like the "concrete distribution" (https://arxiv.org/abs/1611.00712) can help us continuously approximate discrete probabilistic reasoning. Maybe as these techniques move into the mainstream in systems like Venture or Picture we can start to scale up predictive/generative/probabilistic modelling to match optimization-based connectionist methods?
My lay person's impression is that at its most basic level a recurrent neural network is simply a "conveyor belt" of neural nets which are affected by external weights as well as by the weights from within the network. More precisely the "internal" weights coming from the layer of perceptrons operating at 1 level shallower than itself. So we're dealing in essence with 2 dimensions (shallower to deeper, and older to newer) instead of just one (shallower to deeper).
I think the shift in thinking should rather be: instead of trying to build the best possible associative memory to associate some A with some B, take the memory modules we have (perhaps not perfect) and try to build something bigger out of them. A dynamical model of the observed reality seems like a great thing to build out of such modules.
And this is what the PVM is. Currently made out of shallow, plain vanilla perceptrons, builds a structure which can be arbitrarily deep. Without any "magical" tricks such as dropout, relu, convolution, pooling etc.
Could this give rise to self perception or consiousness?
a) Self-perception does not seem like consciousness to me at all. In meditation, if done properly, there is very little self left. It feels more like pure awareness. It is almost the opposite of the model of the self that the brain constructs.
b) I fail to see how the fact that a mechanism refers to itself should somehow give rise to the feeling of conciousness. Why would it? Nobody would predict consciousness from that if we would not already know it exists and it's easy to imagine a device that has a model of itself and is not self-aware.
I understand the need to somehow fit this into our scientific framework, and the idea that "consciousness is just what it feels to have a brain" is the best thing we have, but I don't think it explains anything. There is something we are missing.
Many definitions can coexist, some more actionable than others. "Being aware of the existence of oneself in the world, and being able to reflect on oneself's decision" seems relatively practical. So, Self-perception + self-reflection = consciouness (as a definition)
From this starting point, it seems reasonable to derive that consciousness can arise from 1) mental representation of the world that include oneself 2) empathy for others (I can guess why this other worker has taken this decision) that, once applied to the actions of the self as if it were an external agent, gives self-reflection.
https://plato.stanford.edu/entries/consciousness-representat...
https://plato.stanford.edu/entries/consciousness-higher/
https://mitpress.mit.edu/books/self-representational-approac...
http://www.nyu.edu/gsas/dept/philo/courses/consciousness05/L...
So consciousness was not raised at this point. But that doesn't mean that it couldn't be an emergent property.
[1] Am at NIPS and attended the speech.
I thought this ommission was deliberate to avoid distracting philisophical ratholes that weren't core to his talk.
Also, if anyone is watching Westworld (spoilers), it seems to come to the same conclusion funnily enough. What finally gives the androids consciousness is some kind of recursive idea of listening to themselves.
0. https://www.amazon.com/Prey-Michael-Crichton/dp/0061703087/r...
I found a pdf version of the book here if you are interested. http://selfdefinition.org/psychology/Julian-Jaynes-Origin-of...
It seems the current state of prediction is only slightly better than the state of image recognition pre multiple level NNs
There might be still a theoretical jump that's needed