Differentiable Neural Computers
deepmind.com
deepmind.com
The feeling among researchers I've spoken to is not that NTMs aren't useful. DeepMind is simply operating on another level. Other researchers don't understand the intuitions behind the architecture well enough to make progress with it. But it seems like DeepMind, and specifically Alex Graves (first author on NTMs and now this), can.
Would you be so kind as to to explain what you mean here ?
Thanks !
Here.. take the 'Axon Hillock' https://en.wikipedia.org/wiki/Axon_hillock code up a function for it, attach it to present day neuron models, make it do something fancy, write a white-paper and kazaam you're operating on another level..
Get it?
Another amazing thing they did was to generate audio by direct synthesis from a neural net, beating all previous benchmarks. If they can make it work in real time, it would be a huge upgrade in our TTS technology.
We're still waiting for the new and improved AlphaGo. I hope they don't bury that project.
It's decoupled yet stores transient local information regarding previous neuron activity.
Another bio-inspired copy-pasta.
Many algorithms are bio-inspired -- good artists borrow, the best steal.
Input = Data
Process = Optimisation to create an automata.
Output = Automata
Computer power means much larger variable spaces can be handled in optimisation problems. NN are a means to prune the variable space during optimisation in a domain unspecific way.
That's not to say that NTMs are bad or uninteresting! They are super cool and I think have huge potential in natural language understanding, reasoning, and planning. However, I do think that DeepMind will have to prove that they can be used to solve some non-trivial task, one that can't be solved much more efficiently with traditional CS methods, before people will join in to their research.
Also, I think there's a possibility that solving non-trivial problems with NTMs may require more computing power than Moore's law has given us so far. In the same way that NNs didn't really take off until GPU implementations became available, we may have to wait for the next big hardware breakthrough for NTMs to come into their own.
I think the strength of NTMs will be best demonstrated by putting it to work on a long-range language modeling task where you need to organize what you read so that you can use it to predict better a paragraph or two later. Current language models based on LSTM are not really able to do this.
Nope. It's really easy to solve simple problems; it can sometimes even be done by brute-force.
That's what caused the initial optimism around AI, e.g. the 1950s notion that it would be an interesting summer project for a grad student.
Insights into computational complexity during the 1960s showed that scaling is actually the difficult part. After all, if brute-force were scalable then there'd be no reason to write any other software (even if a more efficient program were required, the brute-forcer could write it for us).
That's why the rapid progress on simple problems, e.g. using Eliza, SHRDLU, General Problem Solver, etc. hasn't been sustained, and why we can't just run those systems on a modern cluster and expect them to tackle realistic problems.
Mainstream work on neural nets is focused on pattern recognition and generation of various forms. I don't mean to trivialize at all when I say this - this gives us a new way to solve problems with computers. It allows us to go beyond the paradigm of hand-built algorithms over bytes in memory.
What DeepMind is exploring with this line of research is whether neural nets can even subsume this older paradigm. Can they learn to induce the kinds of algorithms we're used to writing in our text editors? Given this goal, I think it's better to call problems like sorting "elementary" rather than "toy".
It seems like the way forward would be networking together various kinds of neural networks to achieve complex goals. For example, an NTM specialized in formulating plans that has access to a CNN for image recognition, and so on.
In the human brain, Neurons store an incredible amount of information. Neuron models in neural networks only did so with weights.
There is still a lack of understanding on how the human brain does it. Deep Mind grabbed a proven memory model from Alan Turing's work and applied it to the feature barren neuron models in use. Sprinkle magic ...
They are not operating on another level, they're bringing over features that are well documented in the human brain and in white papers from a past period when people actually thought deeply about this problem and applying it.
https://en.wikipedia.org/wiki/Bio-inspired_computing
There is no 'intuition' about the architecture. Study the human brain and copy pasta into the computing realm.
Others are doing this as well. If anyone bothered to read the white papers people publish, you'll see that many people have presented similar ideas over the years.
You can come up with your own neural Turing machine. Take a featureless neuron model, slap a memory module on it and you have a neural turing machine.
Graves and co. have been really creative in overcoming problems in their ongoing program to differentiate ALL the things.
To train a neural network, you want to know how much each component contributed to an error. We do that by propagating the error through each component in reverse, using the partial derivatives of the corresponding function.
Let's suppose on the current training input, the network produces some output that is a little wrong. It produced this output by reading a value v at location x of memory.
In other words, output = v = mem[x]
It could be wrong because the value in memory should have been something else. In this case, you can propagate the gradient backwards. Whatever the error was at the output, is also the error at this memory location.
Or it could be wrong because it read from the wrong memory location. Now you're a bit dead in the water. You have some memory address x, and you want to take the derivative of v with respect to x. But x is this sort of thing that jumps discretely (just as an integer memory address does). You can't wiggle x to see what effect it has on v, which means that you don't know which direction x should move in in order to reduce the error.
So (at least in the 2014 paper, ignoring the content-addressed memory), memory accesses don't look like v = mem[x]. They look like v = sum_i(a_i * mem[i]). Any time you read from memory, you're actually reading all the memory, and taking a weighted sum of the memory values. And now you can take derivatives with respect to that weighting.
To me, the question this raises is, what right do we have to call this a Turing machine. This is a very strong departure from Turing machines and digital computers.
As for "digital" computers remember they are built out of noisy physical systems. Any bit in the CPU is actually a range of voltages that we squash into the abstract concept of binary.
We don't squash the range of voltages. The digital component that interprets that voltage does the squashing. And we design it that way purposefully. https://en.wikipedia.org/wiki/Static_discipline
Turing specified that the reads and the writes are done by heads, which touch a single tape position. You can have multiple (finitely many) tapes and heads, without leaving the class of "Turing machine". But nothing like blending symbols from adjacent locations on the tape, or requiring non-local access to the tape.
http://www.nature.com/nature/journal/vaop/ncurrent/full/natu...
If AGI is the goal and machine learning research is the search algorithm, then Schmidhuber's attempting to perform backpropagation by pushing rewards back along the connections :)
Maybe I just feel a bit uneasy with a claim such as:
> We hope DNCs provide a new metaphor for cognitive science and neuroscience.
Alex Graves and some others in DeepMind have focused a lot in the past year or so on developing practical differentiable data structures, so that the LSTM can read and write to an external memory (and save its precious internal state for more immediate needs) yet still be trainable via backpropagation.
LSTMs are designed to capture long range dependencies, e.g., "this word at the start of the sentence interacts with this word at the end of the sentence."
DNCs are designed to incorporate outside information, e.g., "i happen to know (from background knowledge) that these two people in this sentence are married"
Present day Neuron models lack an incredible number of functional features that are clearly present in the human brain.
NTMs = representing memory that is stored in neurons https://en.wikipedia.org/wiki/Neuronal_memory_allocation
Decoupled Neural Interfaces using Synthetic Gradients = https://en.wikipedia.org/wiki/Electrochemical_gradient
Differentiable Neural Computers = Won't specify what natural aspect of the brain this derives from.
Pick an aspect of a neuron or the brain that isn't modeled, write a model...
Bleeding edge + Operating on another level
The fact that someone is going out of there way to remove points from my posts so that this doesn't see tomorrow's foot traffic instead of replying and critiquing me just goes to show how truthful these statements are.
Anyone can create such models. No one has a monopoly or patent on how the brain functions. Thus, expect many models and approaches.. Some better than others.
You can down-vote all you want. The better model and architecture wins this game. It would help the community if people were honest about what's going on here but people instead want to believe in magic and subscribe to the idea that only a specific group of people are writing biologically inspired software and are capable authoring a model of what is clearly documented in the human brain. Interesting that this is the reception.
"Neural Turing machines" are not the same thing as neuronal memory allocation: NTMs' memory is external and neuronal memory allocation is all about how memory is stored in neurons in the brain.
The "synthetic gradients" in that paper have nothing to do with the electrochemical gradients you mention other than the name.
No one is claiming that the DeepMind guys are "operating on another level" because they do bio-inspired things. They are claiming that because they are getting more impressive results than anyone else.
Now: Are they really? If so, is that enough justification for such a grand-sounding claim. I don't know. That would be an interesting discussion to have. But "Boooo, these people are just copying things present in the brain, there's nothing impressive about that" is not, especially when the parallels between the brain-things and the DeepMind-things are as feeble as in your examples.
Making statements that allow people to see behind the curtains and maybe go off and make their own competitive models... Yes, this is a disservice to the advancement of A.I and should be downvoted : Removing the prestigious veil and illusion from published works.
NTMs memory is external in what sense? Please detail what this means in a 'functional' sense. It's biologically inspired. Neurons maintain memory beyond synaptic weights. The neuron models of present day A.I were basic. Someone comes along and sees the obvious : There is no computational model for how neurons utilize memory and suddenly they're thinking on another level? Give me a break..
Synthetic gradients have everything to do w/ electro-chemical gradients : http://www.nature.com/articles/srep14527 http://www.pnas.org/content/110/30/12456.full.pdf So, where is your establishment that I am incorrect. It is nowhere to be found. Again, biologically inspired computational models.
Oh look, someone published a paper back in June that is an implementation of Differentiable Neural Computers: https://arxiv.org/abs/1607.00036
It's hype and that is a disservice to the community of people completing similar work and taking similar approaches.
It would be an interesting discussion to have. That discussion was terminated in favor of downvoting me.
They're feeble to someone who isn't well informed on neuroscience. Thus, you'd rather be wow'd and believe in the fantasy that only a small segment of people can write computational models of biology.
Continue believing the hype. Rarely will someone be truthful and honest about where they got their ideas when hype follows. An interesting conversation could have transpired. Enjoy the feels from the downvotes.
They are the first who have a machine learn to solve problems that require memory. They are the first. These are the stepping stones to artificial Intelligence.
Note: The whole point of the Synthetic gradients, is to learn a network in parallel. This allows Google to make computers learn recognize things in images even better. To recognize human speech even beter... To make self driving cars even better.....
I don't know if they are copied or not from nature (doesnt look like). The point is that they are improving mankind.
Incorrect. It was named a Neural (Turing) machine for a reason. Maybe people should go back and dust off the white papers from the 70s like those who are borrowing from that era and respectfully giving credit where credit is due.
They do great work and they are making great progress in Artificial Intelligence. Many people are. Everything is a stepping stone. It serves no good to over-hype one person's stones over another's or ignore/downplay where they were inspired from. Notable visionaries of a past time were visionaries because they detailed the depths of their thinking and centered on the hows/whys. It seems it is fashionable now-a-days to do the exact opposite. This is to a disservice to learning and progress.
The whole point of the human brain is parallel processing. Extra-cellular chemical Gradients function the same way in the human brain and serve the same purposes. Take a look at the papers I linked.
> I don't know if they are copied or not from nature (doesnt look like).
Extra-cellular chemical Gradients. I linked to white papers that explain how memory is stored in them and shared across neurons. This is how it works in nature and biology.
They named their approach 'Synthetic Gradients'. An artificial form of the biological Gradient that is decoupled and lies outside of a neuron. They are clearly giving credit to nature.
They and many other people are improving mankind. Many others can improve mankind if there was less hype and more of a focus on where the ideas originated.
That was my point..
The behavior of people regarding selective 'hype' is one of the big reasons why a tremendous amount of deeply functional work that centers on hard intuitions and ideas for this area will remain closed source when a real break is made.
Enjoy the hype train I guess... They're operating on another level than anyone else.
An Synthetic Gradient is a way to allow learning Forward Propagated Neural Nets in a parallel way. The gradient here is referring to the 'error' backpropagation that is part of the training process of an neural net (im talking about computer science neural nets).
They have nothing todo with each other. The papers that you are referring to have nothing todo with the process of training a neural net.
Coming up with a functional systems architecture that ties the bits and pieces together is hard work. Understanding what is really happening in the human brain, how/why it is performing various functions, and how this provides for an intelligent architecture is hard work. Creating an 'aware' platform is hard and elusive work which is why people chase the low hanging fruit of optimization algorithms.
*Cheers