At best this research says something about the task of navigation and optimal representations for that task rather than anything profound about neural networks other than they can both optimize for some task—which should surprise no one.
477 karma · joined August 22, 2011
twitter.com/@jimmfleming
At best this research says something about the task of navigation and optimal representations for that task rather than anything profound about neural networks other than they can both optimize for some task—which should surprise no one.
Points are processed as local clusters. Each cluster of points is ordered according to a transformation matrix X. This produces a "canonical" ordering so the processing of points in these local clusters do not need to be invariant to the order since the order now has consistency and meaning.
It's kind of like placing each point into a regular grid like an image before running the convolution. The trick is which points to put in which grid cell which is determined by the transformation matrix X. This takes advantage of locality which PointNet does not, if I recall correctly. By acting locally and stacking many layers of these you can produce a hierarchy of more and more abstract clusters of points, each with an inherent relationship to nearby clusters of points. In addition, the transformation matrix also appears to act as an attention over the points in the cluster.
[0] http://gamma.cs.unc.edu/RVO/ [1] http://gamma.cs.unc.edu/RVO2/ [2] http://gamma.cs.unc.edu/HRVO/
Generalization is a multi-axis scale, not a switch: you can have more or less generalization in many different dimensions. Being terrible at adversarial examples just means that axis is weak.
Memorization is pretty easy. Generalizing from past examples requires that there be a relationship not just between one person and their phone number but between all people and their phone numbers.
A better way to summarize the central question of this paper would be: "Why is it that a large-parameter model trained with gradient descent on real data _could_ just memorize all of the training data (it has the capacity) yet finds solutions which generalize well to an unseen test set?"
To say that deep learning is _just_ memorizing its training data would be incorrect. We have empirical evidence to the contrary and this paper is part of that evidence.
> Both RNNs and CNNs can stabilize learning via weight sharing, therefore they are less prone to these perturbations. In contrast, FNNs trained with normalization techniques suffer from these perturbations and have high variance in the training error (see Figure 1).
Essentially FNNs stand to benefit more from this work than CNNs or RNNs.
Graphs have advantages but they can be unfamiliar and sometimes difficult:
Some advantages: easier to serialize the whole graph and distribute computation; optimization can also be performed across the graph nodes (e.g. see XLA in TensorFlow).
Some disadvantages: it can be more difficult to write and reason about, particularly for recurrent neural networks which can utilize loops a lot; also interop with reinforcement learning environments where much of the computation is performed in an environment outside of the graph.
"Towards a Biologically Plausible Backprop" from Benjamin Scellier and Yoshua Bengio (2016) would be a recent paper on the topic.
In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.
My reading of [2] is that the model only needs to be retrained if the environmental boundary changes. For example, to handle objects that do not fit within the current boundary. Since I imagine scale matters here, and you couldn't simply make everything inside smaller. Internal geometry such as replacing the castle with a car does not appear to require retraining.
The point is that there will always be haters, so being discouraged by them or complaining isn't fruitful.
I am surprised by #2: "Build a household robot". It's my understanding that efficient actuation and power are largely unsolved problems outside of the software realm. What's the plan for tackling stairs, variable height targets, manipulator dexterity, power supply, etc. in a general purpose robot with off-the-shelf parts? (Answering these questions may be part of that goal but maybe someone knows more on the subject.)
From the Hodgkin-Huxley (HH) neuron model to integrate and fire (I&F) you can approximate and simulate various levels of realism in artificial neural networks. There's even the Nengo library[0] which can simulate these for you in Python. Unfortunately, they're very hard to work with compared to perceptrons, artificial neural networks, and deep learning (see Neural Engineering[1] for a good framework). Additionally, they're far more computationally intensive with the HH model coming in at over 800 FLOPS to the I&F model (which lacks a lot of the behavior and dynamics of biological neurons) which has 7 FLOPS. Then try representing numbers with spiking neurons... You need several dozen, working as a population[2], to accurately represent a single floating point because they're so noisy. This makes working with data pretty difficult and expensive.
[0] https://github.com/nengo/nengo
I'm not sure what this means or how the abstractions are meaningless? From Gabor filters to concepts like "dog", the abstractions are quite meaningful (in that they function well), even if not to us.
> They are not. Hundreds of man years worth of engineering time go into each of those systems, and none of those systems generalizes to anything other than the task it was created for. That's nothing like human intelligence.
This isn't strictly true if we look at the ability to generalize as a sliding scale. The level of generalization has actually increased significantly from expert systems to machine learning to deep learning. We have not reached human levels of generalization but we are approaching.
Consider that DL can identify objects, people, animals in unique photos never seen before and that more generally the success of modern machine learning is it's ability generalize from training to test time rather than hand engineering for each new case. Newer work is even able to learn from just a few examples[0] and then generalize beyond that. Or the Atari work from DeepMind that can generalize to dozens/hundreds of games. None of those networks are created specifically for Break Out or Pong.
It's also not entirely fair to discount the hundreds of years of engineering considering most of these systems are trained from scratch (randomness). Humans, however, benefit from the preceding evolution which has a time scale that far exceeds any human engineering effort. :)
It has similarities to backprop, and depends on backpropagated signals but backprop is simpler, faster and can exploit knowledge about the loss function (as far as we know STDP cannot). The downside (?) to backprop is that it cannot exploit temporal information like STDP but several recurrent models have found ways around that.
[0] https://www.allegorithmic.com/products/substance-painter [1] http://quixel.se/
1. The pace of published research is pretty fast right now. This makes it difficult to know where the research fits in when solving problems. It'll probably take a few years before we know where to use many of the approaches published last year.
2. Iteration performance (trying many new things quickly) is improving with high-level frameworks but still lots of work to do here. Since we don't know where a lot of research fits it's not always apparent which methods work best (and "best" changes every quarter).
3. We're still missing theoretical foundations for much of deep learning. This is useful, not just for research, but to know what can and cannot work with current approaches.
4. Model architectures are still based largely on trial-and-error, intuition and search.
The article describes the outcome of Lee Sedol's match correctly:
> Lee went on to lose all but one of their match’s five games.