How Google understands language like a 10-year-old
sfgate.com
sfgate.com
I have a more specific question about Google Search that I'd like to see answered. To what extent do they model specific languages, versus training classifiers? Are they really grokking sentences or sentence fragments, or do they have enough training data to fake it, like Bill Gates in "Petals Around the Rose" [1]?
[0] http://googleblog.blogspot.com/2010/01/helping-computers-und...
The problem is that natural languages are at least context-free, and no algorithm exists (can exist?) to parse context-free languages in less than polynomial time, with respect to the length of the sentence. You can approximate by parsing with probabilistic finite-state machines, but they get led down blind alleys and can't backtrack, so they're inaccurate. Let me know if you'd like me to elaborate on this with examples (or you can Wiki "garden path sentence" and probably imagine the problem).
I'm also sure they're not doing supervised word sense disambiguation. That's in a poor state too, and imho isn't even a good idea. The whole concept of having someone list out the "senses" of a word is misguided, because it's totally unclear how fine-grained you should be. And then you need at least a couple of hundred labelled examples for every word...
Most of the examples they give are best explained by dimensionality reduction techniques, which have been popular in information retrieval for some time. Google have undoubtedly invented some secret sauce, but they've also just got orders of magnitude more data and processing power.
But yeah, context sensitive stuff can be way worse.
Unless new evidence has surfaced in the years since I finished a degree in linguistics, there's only a couple of pieces of evidence for language constructions in natural languages that can't be generated by a context free grammar, like a Adv1Adv2Adv3Adj1Adj2Adj3 construction in Zürich dialectical German (where Adv = adverb and Adj=adjective, and numbers represent which adverb modifies which adjective).
I do have a long history of over-analysing simple mathematical games, though.
Anyway, these days I'd obviously be wrong. I'm so impressed with the strides Google has taken to NLP, and I am fully expecting them to beat everyone to Strong AI. And why not? They know that the better they are, the better their advertising revenue will be. And they know that once they get there, even if ads are no longer profitable having the world's only AI will be incomprehensibly popular.
My one problem with this article is the last line:
"They're still not approaching the conversations you'd have as a teenager."
Google hasn't yet approached "the conversations" you'd have with a 5 year old. While Google may understand a 5-year old's conversation, it certainly couldn't participate it in and reply back to the kid.
I would argue that it sort of does. Only instead of a normal kid, it's a mute kid that can only reply to you by passing you back documents it thinks you're asking for.
http://www.google.com/search?q=what+is+the+height+of+the+emp...
http://www.google.com/search?q=what+is+the+boiling+point+of+...
"1,250".
What?! 1,250...feet? inches? meters? centimeters?
Are you sure it's only the technological advance that's responsible for that? I've always written search terms like that because I figured the search engine isn't some ai that's answering my question. Instead it's a program searching a database of sorts and by searching for what I think other might have searched for I'm able to get similar results.
Google's understanding of language is similar to that of a savant who has been imprisoned since birth and has been tied to a bench in front of a screen showing texts of the web. He can't read in the sense that he could pronounce words, but he recognizes familiar patterns of symbols. He does not know what "pancakes" are, but he knows that the word is often seen with the word "butter". It's amazing how much can be done in this way, but it is quite different from how humans understand.
Nope. As I read the article, I thought "Cool! it has trouble with negation, just like my 2 year old." A few months ago, it was obvious that if I said "Don't do X", she'd just match on "X", and do the thing. E.g. "Don't poke your eyes" => pokes eyes.
She's starting to get the hang of it now, but I can see how confusing it is for her.
EDIT>
Also, babies can understand a lot before they're able to compose language. E.g. "bring me your shoes" ... brings shoes. "Give that to Mommy." ... gives thing to Mom.
It's all about the FedEx quests, initially.
My son told his first joke before he could form sentences. He was about 10 months old; I was getting him dressed, lying on the changing table. "OK, give me your hand." He lifted up one hand so I could put it into the sleeve. "And now your other hand." He got this sly grin and lifted his foot.
I used to believe, like you do, that the ability to parse and make decisions based on input was not the same thing as understanding.
Then, I wrote a chess bot for a CS lab. The thing plays better than me, better than it's peers. My partner, who was good at chess (or at least very literate in it) could identify what strategies it was going for. We had a visualization of what moves it was considering, and you could see that it was essentially playing chess by swinging a baseball bat around and seeing what looked nice.
Does the chessbot "understand" chess? It sure seems like it. We like to think humans are special, and have some kind of unique understanding that computers can't, but I think it's only us lying to ourselves.
People always seem to look at google as though they're doing some special secret magic, but in fact they're really just implementing fairly well-known CS algorithms. They just do it exceptionally well.
Also a paper from someone at Google on brute force paraphrase acquisition using billions of sentences combined with some relatively simple rules.
http://www.scribd.com/doc/13863110/The-Unreasonable-Effectiv...
That said, I don't think tagging for them would be very simple like you say. For a start they're dealing with multiple languages, probably many languages without any human annotated training corpora. Even for the languages with training data, web pages are difficult to tag & parse because they often contain very 'slack' grammar and domain specific/slang words. The standard English training corpus is the Penn Treebank (Wall Street Journal text), can you imagine trying to read and understand youtube comments if all you'd ever read was the WSJ? Even tagging search queries would be difficult because they're not even sentence fragments you could use viterbi with, they're often just words strung together without any grammar construct at all so you can't rely on the tag order you know from your corpus to help you tag a query.
So I'm very impressed that they're doing any tagging at all, on the scale they're doing it at, and with presumably decent enough results for it to be useful.
But still, all the obvious ways to use POS tagging for IR have been tried, and don't work. POS is only just marginally useful in machine translation at the moment!
This isn't that surprising when you think about it. What POS tags provide is a small clue about grammatical structure. So you're either setting off down that path and trying to understand a sentence as a tree, or you're going to understand a sentence as a sequence of words. Syntax is hard, and most word-word dependencies are between adjacent words.
There's a local maximum at "just use the words", and POS tags don't take you far enough to get past it for many tasks.
I agree, the bag of words / vector space model approach is (from what I've read) hard to beat, especially for a one-size-fits-all tool like google where you can assume nothing about the domain. (Unless they're switching approaches depending on the domain, which would be fascinating)
I suspect we have some miscommunication here.
Experimented with the Brill POS tagger, which I think is a little old now.
I can't speak for the state-of-the-art, but that's certainly what I was taught last year. I studied both computational linguistics and IR -two very different approaches to the same problem. IR seemed to be getting much further in terms of real-world results than linguistic approaches. Although linguistic approaches may seem to offer greater potential.
I was delighted the day I noticed it knew that "regexp" and "regex" were synonyms for "regular expression".
If this is a linear learning curve, in another 15 years it should just start to be able to 'understand' the nuances of Shakespeare and Ulysses among others.