Google Is All About Large Amounts of Data
googlesystem.blogspot.com
googlesystem.blogspot.com
Thats how google is attacking the "parsing text problem" to find meaning. [0] Not with regular expressions, rules or clever AI hacks. Just plain old math.
[0] Attributed to Peter Norvig. Here's an example of what is suggested. A spell checker in about 25 lines python (old but good) ~ http://norvig.com/spell-correct.html
21 lines in Python 2.5 code
Its a good talk, though maybe should be called "theorizing from massive amounts of data". Makes you wonder what Google are keeping under wraps for now, and how the powerset approach can compete, except maybe for domain-specific stuff.
google "knows" (believes with high probability, I guess) that PWC is an abbreviation of PriceWaterhouseCoopers, for instance.
you can see this directly in how google highlights terms in search results. you can certainly find instances where they "should" have gotten something but didn't, or made a mistake, but it works well enough most of the time.