Alibaba neural network defeats human in global reading test
zdnet.com
zdnet.com
The answer to every question in the test is a preexisting snippet of text, or "span," from a corresponding reading passage shown to the model. The model has only to select which span in the reading passage gives the best answer -- i.e., which sequence of words already in the text best answers the question.[a]
Actual current results:
https://rajpurkar.github.io/SQuAD-explorer/
Paper describing the dataset and test:
https://arxiv.org/abs/1606.05250
[a] If this explanation isn't entirely clear to you, it might help to think of the problem as a challenging classification task in which the number of possible classes for each question is equal to the number of possible spans in the corresponding reading passage.
Compare: "ROBOTS CAN NOW READ BETTER THAN HUMANS, PUTTING MILLIONS OF JOBS AT RISK" http://www.newsweek.com/robots-can-now-read-better-humans-pu...
Before you blink an eye there will be some MBA-types working on PowerPoint proposals with detailed cost-benefit analyses for using those new AI machines they heard about that can read better than human beings. Needless to say, the technology will fall far short of expectations.
This is why there have been two AI winters already.
The futurist writers peddling this stuff need to take a moment to chill and learn about the actual state of the underlying technology.
Bloomberg news: "Alibaba says it’s the first time a machine outperformed people" "China’s Plan for World Domination in AI"
Sounds like they've reinvented Jeopardy Watson's ability to excel at Q&A, but 12 years later.
That said, I think the path to 'real' AGI lies in some combination of DL, probabilistic graph models, symbolic systems, and something we have not even imagined yet. BTW, a good paper just released on the limitations of DL by Judea Pearl https://arxiv.org/abs/1801.04016
Well, that really made it much clearer to me ;)
> Mismatch occurs mostly due to inclusion/exclusion of non-essential phrases (e.g., monsoon trough versus movement of the monsoon trough) rather than fundamental disagreements about the answer.
I don't think I would call that "error," rather than ambiguity. In other words, there's more than one possible answer to the questions under these criteria -- English isn't a formal grammar where there's always one and only one answer. For instance, here's one of the questions from the ABC Wikipedia page:
> What kind of network was ABC when it first began?
> Ground Truth Answers: radio network radio radio network
> Prediction: October 12, 1943
Because the second human said "radio" instead of "radio network," I believe this would count as a human miss. But the answer is factually correct. Meanwhile, the prediction from the Stanford logistic regression (not the more sophisticated Alibaba model in the article, where I don't think results are published at this detail) is completely wrong. No human could make that mistake. And yet these are treated as equally flawed answers by the EM metric.
And yet this gets headlined as "defeats humans," not "learns to mimic human responses well."
For some problems, sure. For prediction tasks on the other hand, you have an actual ground truth that can be compared to human a priori prediction.
Neural net NLP results are rarely about actual intelligence or clever use of latent variables that it figured out, and more "pattern matching" that explains why its errors are so different from human errors. It doesn't actually understand the problem, it's finding tricks to solve the questions that happen to be regularities in the dataset that us humans can't really see.
Nevertheless, most extractive systems learn some degree of co-reference resolution.
I have a less advanced system than the Alibaba one, and it got both example questions correct:
The trophy would not fit in the brown suitcase because it was too big. What was too big?
and
The town councilors refused to give the demonstrators a permit because they feared violence. Who feared violence?
Also, the Alibaba Cloud is looking for engineers. Pls check https://careers.alibaba.com/positionDetail.htm?positionId=b7...
Did someone change the link?