Solving Verbal Questions in IQ Test by Knowledge-Powered Word Embedding [pdf]
arxiv.org
arxiv.org
This is a bit like saying, "the best runners tend to be pretty tall, and we've made a robot which is really tall – so running robots are just around the corner."
IQ tests certainly correlate well with intelligence (in humans), but they're a metric, not the thing itself. Another metric would be mental arithmetic; people who can do sums quickly tend to be pretty smart, but that doesn't mean that calculators are a step away from super-intelligences.
Interesting and cool work but let's be careful in the interpretation (and remember that people said the same things about chess playing not that long ago).
That being said, making an AI that does better on IQ tests than humans is a rather interesting and worthwhile endeavor.
The decades old system for handwriting digits recignition for US Post beated humans (now it is a few lines of Octave in Andrew Ng's course). Still it cannot write a reply to a letter.)
I understand this isn't intelligence and the title doesn't imply it but I'm not sure what your reference to decoding is about.
The test they use to see if it actually learned what these words meant, in a limited sense, is to test it against a subset of verbal IQ tests (not what it was trained on!). You could ask it the antonym, synonym, or analogy for anything in English. This is an extension of word2vec / word embeddings.
That it beats the scores of college graduates impresses me.
I don't think that is entirely correct. After cursory reading of the paper, my understanding is that they look up a list of word senses for each word in a dictionary (or multiple dictionaries). And then they try to learn something about each of those word senses from wikipedia (that is they create seperate word embeddings for each of those senses). So what they do not do is to learn what senses a word has. That is done by the humans who created the dictionaries.
What that means is that they cannot pick up new senses of words, which doesn't matter for answering IQ test questions because these questions rarely change and are typically based on well established word meanings.
Unfortunately it makes this approach less than ideal for things like understanding the news (something I'm working on), where new contexts of words keep popping up all the time.
See here for an MIT experiment in blurry text transcription: http://groups.csail.mit.edu/uid/deneme/?p=329 for the unexpected accuracy resulting from crowd-sourcing.
And generally humans don't natively know when they're experiencing an optical illusion. They have to be taught it. And in either case, it's not the human vision system that learns the lesson, it's some other part of the brain that learns to discount the vision system's conclusions.
... wait, don't we already do that?
One thing that popped out at me, however, was the distribution of the human scores: monotonically increasing with age, which was somewhat odd. Shouldn't it be more or less normally distributed?
1. I wouldn't have thought to make a serious attempt towards beating a verbal comprehension test with deep learning; the possibility of succeeding would've seemed tiny in compared to the work. Similar to the notion that trying to prove the Riemann Hypothesis is hard to think of as a real pursuit, at best it's like a quixotic hobby.