A Theory of Sequence Memory in Neocortex
blog.acolyer.org
blog.acolyer.org
This is distinctly different than say a recurrent neural network where all neurons have outputs at each timestep and weights are updated based on the derived contribution of each weight to the quality of a final output.
The first rule is very local, both in time and space, and the second is in some senses "global."
Backprop has obviously had much more success in ML than any Hebbian learning rule. However we're pretty sure biological learning is essentially Hebbian. The merger / reconciliation of both these rules is one of the most interesting areas of research right now.
Edit: linked to specific post in forum thread instead of the whole thread
I’m drawing a blank how you’d run a nn that way?
Most of what we're seeing now with neural networks was pretty well understood 20+ years ago. The hard part is forming hierarchies or clusters of networks that are only active in certain scenarios, so that we can scale learning beyond simple pattern recognition (which is effectively a solved problem today).
I've read very few satisfactory articles on how to go about teaching large neural nets to learn multiple distinct patterns and know which networks to engage in different scenarios. This article is suggesting that dendrites form that orchestration layer by waiting for enough specific patterns to happen in close enough proximity that the neuron becomes engaged and goes to work when the rest of the pattern arrives. This gives the network some ability to predict what is coming, whereas the simple neural networks we model would only fire when the whole pattern matches.
There is a bit of hand waving here and I probably missed something so please add any insights of your own thanks!
I am very curious to know what labs/companies are using this approach.
See this blog post for why I chose consciousness as the primitive to build (instead of say demoing the model's characteristics on common tasks like mnist digit recognition) - https://medium.com/creating-artificial-consciousness/the-cas...
Does anyone understand why this is true?
In a typical multi-layer network, doesn't each node in a lower layer connect to every node in a higher layer? All the {L(i, n-1), L(a, n)} edges going from the nodes in layer n-1 to to a particular node (a) in layer n would constitute a dendrite.
The way some authors explain it is that you can think of a neuron as really having an entire neural net embedded inside of it (based on the structure and function of the dendritic tree), that does some non-trivial amount of information processing even before you consider the connectivity of neurons to each other. Exactly how non-trivial that is is a deeper question, but it's worth noting that these structures are extremely plastic, and change over timescales of seconds to minutes, so it's not hard to imagine that these details are significant.
I also read Hawkin's book and loved it. I'm still hoping that they'll get something going.
Prediction is only useful for an organism inasmuch as it allows it to calculate expected future utility for actions. And these actions are the only things that lead to different outcomes in terms of evolutionary fitness, thus the actions that are ultimately output by the brain are the only things that determined its evolution.
If the purpose of prediction is to calculate expected future utility of different actions, then it does not follow that a general-purpose prediction device will be useful, because the prediction device might use all its energy predicting aspects of the environment along dimensions that are irrelevant to utility. A useful prediction device would only predict along useful dimensions, and may be very different in behavior from a generalized autoencoder. As an example of this difference as it comes up in the field of deep learning, consider the case of machine translation: you could either train a sequence model (LSTM or whatever) to autoencode sequences---i.e., to be able to predict them---and then use the resulting representations to do translation, or you could train end-to-end where the objective function is translation quality. It turns out the latter yields better results.
Maybe the brain, or part of the brain, is a prediction engine and another part does action selection based on the predictions. But then why would we identify intelligence with the prediction engine part rather than the two parts combined? Searching for an optimal action is a very different task than prediction; it seems you need both to have what anyone would call intelligence.
I'll try to summarize (my own layman's understanding of) the Predictive Processing theory further here, for people who don't want to read the article:
As the PP theory goes, not only is the human brain a "general-purpose prediction device", but it has no other components. When we think, we're simply making predictions for the way the world will be a moment from now; and then we attempt to reconcile those predictions with the image we get of the world a moment later. In equal parts, depending on our confidence in our senses vs. our confidence in our model, we either do this by "learning" (i.e. correcting the predictions to match the signal) or by "acting" (i.e. correcting the signal to match the predictions.)
The latter, "acting", is done by simply playing out the error signal (i.e. the difference between "the way the world is" and "the way the world should've been") to motor neurons, which interpret those error signals as motor commands. We seek to be less confused by the world by forcing the world to become more like our model of it!
And (terminal) preferences—those are just persistent exogenous-to-the-predictive-process chemical messengers that bias neurons into outputting a different "world that should've been" than a disinterested observer, only interested in predicting, would've generated. This causes the brain's "acting" process (i.e. correcting error by changing the world) to change the world in the direction of one's preferences, rather than in the direction of one's model. In effect, brains minimize error between the world that they observe, and the world they'd like to observe.
PP theory doesn't say much about where those exogenous chemicals come from, but obviously the neuroendocrine system exists, and consists of a bunch of glands that throw chemicals at the brain.
On the surface he seems to be cramming a whole lot of "solutions" into his HTM (hierarchical temporal memory). HTM is an interesting implementation and the sparse coding is definitely a benefit. However I think he is focused too much on his baby and not on other techniques that may fulfill the necessary components more efficiently.
That is just on the surface. With a product / research balance maybe we just aren't seeing all the cool things going on underneath in research, but it does seem like that research will be shoe-horned into HTM whether or not it is the best architecture.
[0] https://www.youtube.com/watch?v=4y43qwS8fl4&app=desktop
Starting at ~8m20s
This part is so vague. It seems to lack an explanation of how interneurons inhibit other neurons nearby. Also, wouldn‘t sparsity even occur without the early firing enabled by distal pattern matching?
> When relatively few neurons are active relative to the population, then such pattern recognition is robust.
Why?
https://www.biorxiv.org/content/early/2017/10/03/197608
As far as sparsity supporting robust pattern recognition, this paper details the math that shows this:
At a high level, it seems that the constraint of having only a few neurons active is equivalent to a simplicity constraint, thus implementing Occam's Razor and yielding generalization.