Are There Deep Reasons Underlying the Pathologies of Deep Learning Algorithms? [pdf]
goertzel.org
goertzel.org
For those interested, the papers and discussion are part of this weekly collection of AI-related news and resources: https://aiweekly.curated.co
The author is from the OpenCog group, who have spent years building a structured knowledge base as a predecessor to building an artificial general intelligence. It isn't clear how AGI emerges from this work.
The OpenCog KB is a useful piece of work, but it's interesting to note that word/phrase embedding models (word2vec etc) can give similar or better results on most practical tasks that you'd use OpenCog for.
My view is that Deep Learning techniques are insufficient for an AGI, but are good candidates for component parts in the same way that the human optical system does significant preprocessing before hitting the "intelligent" brain.
Also thing like memory networks specifically address some of the episodic memory issues the author raised.
I'd note that Stanford's DeepDive explicitly claims to be a KB building tool, despite it doing probabilistic inference. In their case they do Gibbs sampling instead of PSL, but the similarities are clear.
Also, people associated with OpenCog (including the author of this paper) often refer to OpenCog knowledge bases. Eg https://opencog.wordpress.com/2008/09/14/progress-update/
That's probably correct. The author discusses a known problem; those mis-labeled images are well known, and have been discussed on HN before. It's clear that feature extraction in deep learning sometimes fastens on irrelevant features that somehow work. Some algorithms result in models where data points are too close to at least one decision boundary in a high-dimensional space, which makes them brittle when faced with small, noise-type changes. I don't know enough about the subject to know how that will be fixed, but there are people working the problem and they don't seem to be stuck.
Then the author, who is from OpenCog, goes on to claim, without supporting evidence, that OpenCog can somehow fix the problem. The paper proposed an "internal image grammar", but doesn't say much about what they mean by that or how to do it. Trying to decompose images into some symbolic representation has a long, disappointing history. The computational neural network people are getting results without doing that.
I think we're reading a grant proposal here.
I know some people working using everyday knowledge to attempt to do better image labeling. The idea is that a tree is much more likely to appear in a park than in a kitchen, so you can bias the probable interpretations using that.
I guess that's what they could be talking about. But you need a CNN to get the basic partial label in the first place.
This would also fit well into memory networks: http://arxiv.org/abs/1410.3916, SDM: http://en.wikipedia.org/wiki/Sparse_distributed_memory or global workspace model: http://en.wikipedia.org/wiki/Global_Workspace_Theory
How is this a real paper ?
The author is stating a hypothesis but no way to test it. I'm not sure what point he's trying to make.
Then once you got this dictionary of image "words", next thing is to infer how these words interact with each other, i.e. build a grammar (it's also called "grammar induction" in natural language processing http://techtalks.tv/talks/deep-learning-of-recursive-structu...).
By learning a grammar you essentially define a concept of "chair" or "cat" at a higher level of abstraction by determining how these forms relate to their world (i.e. to other forms), e.g. you can determine that "cat sits on a chair" is a legal phrase in the grammar and "chair sits on a cat" is not.
So extracting a grammar (visual or linguistic) from training data is equivalent to restricting the system to common sense reasoning which operates on concepts in terms of "production rules" of the grammar: http://en.wikipedia.org/wiki/Production_(computer_science), and it is a basic goal of AI.