Deep or Shallow, NLP is breaking out
cacm.acm.org
cacm.acm.org
I keep reading this quote, showcased in the article, again and again, and I still cant believe that people are actually proposing this.
They're basically advocating that we should abandon logic, stop trying to reason about things, and instead try stuff at random, and why? Because computers have been shown to do alright with that approach at tasks in which humans excel.
It is clear as daylight that whatever humans do, computers don't - otherwise, you'd need to train your baby with the Brown Corpus before it could figure out you're its mama. We have managed to overcome the limitations of our primitive computational technology with some clever tricks and that's amazing.
But to take that rightly celebrated fact and make it into an argument that we must now become ourselves as dumb as our computers, and also never try to make the poor things intelligent in the same way we are, that's ... well, it's dumb. That's what it is.
But in "Thinking Fast and Slow", Daniel Kahneman makes that same point that most of our reasoning is kind of heuristic pattern matching. There are lots of experiments that show evidence of this behavior cited in his book. Certainly people can turn it off and logically reason about things when desired, but it's slow and taxing, and the heuristics people live by mostly work, so it's frequently unneeded.
This knowledge is composed from other, simpler and more basic facts, that it slowly learns along the way:
1) you are a reliable source of warmth and comfort
2) you make me feel safe when I'm anxious and scared
3) you feed me when I'm hungry
4) you always come back after going away
5) you make me laugh and feel happy
These can be broken down into more basic facts, as the concepts here are still rather advanced (for a baby). The baby doesn't actually know what the it means to feel "safe", "fear", "happy", etc. It just has fuzzy experiential feelings that slowly become sharper and more contextualized as the baby is "trained" every second of every day by interacting with the world around it.I wouldn't describe them as disparate logical components. The numbered points I wrote are the abstract, rough, adult-level interpretations of the messy "knowledge" that the baby develops. The baby doesn't have this type of vocabulary or even scope of understanding. But I think those would be the flavor of the pillars of knowledge the baby builds up to develop a conception of "mother".
The 2-year old's conception, when they first speak "Mama", will still be incredibly rough. It will be a faint glimmer of what we actually mean when we use the word as fully grown adults. But each additional day more understanding is built up as the network of neurons and their activation patterns, etc is being modified through further interaction. The experiential data from every additional moment of time that passes are pattern matched, integrated, and added to what it already "knows". Humans are really, really good at loose pattern matching & loose analogies. It feels to me as if most of how we think and what we know is just many, many layers of this.
Side point: I would say that even two adults never mean the exact same thing when they use the same word, as they have built an understanding of that word/concept from different life experiences. We can try to communicate to others what we think a certain word/concept means to us, however we have to use other words/concepts to do that, which makes the problem a little difficult! This is why people very often talk right past each other.
You are 100% correct, this is a lot of speculation. Hopefully we will make more progress towards knowing the hard science of how it actually works in our lifetime =)
"You cannot find things by looking for them"
Any person who have ever tried to solve problems (not just rationalize next steps based on previous findings) will know that there is really something unfruitful and irreproducible about logical reasoning.
One of the primary reasons is that logical reasoning you make is based on it's premise. As an example you can rationalize your way through Game of Thrones as long as your premise is set to accept a number of sub-conclusions. Or be a theologist rationalizing your way through the bible.
But this powerful process is also a cage which keeps you trapped inside the same logical consequences of your premises.
And so to invent something new you have to be "irrational" to break out of the boundaries of the premise.
So the claim is not to abandon logic but the belief that problems are solved and inventions made by rational logical process alone. It turns out that most things are actually more trial and error which again also historically has some merit.
Yes, that's called the closed-world assumption. Something is true if you can prove it from the things you know.
You know how else we call the "things you know"? Data.
You know what else we do with data? Train machine learning algorithms.
You know where machine learning algorithms would be if they abandoned their (implicit) closed-world assumption?
The point is that premises allow us to make logical reasoning but the premises themselves are not necessarily logic or universally true or objective or any other notion of right or wrong.
Look through history. Most of the really big innovations weren't a product of some objective but rather incidents and curious phenomena we never expected.
Rational reasoning will take you some of the way as long as you have the premise, but the question is how to get to the premise.
You can't tell whether overfitting is happening by just looking at the input data. You'd at least need a holdout set for testing.
You have to try and remember that our current computer circuits are (boolean) logic, implemented in hardware.
In fact, there's a very good reason why that is so. Turing, and the logicians of Russle's and Goedel's generation before him, conceived of computers as logic machines that reasoned like humans. In Turing's time, computers with logic circuits were even proposed as artificial intelligence. And you know what? For their time, they certainly were.
Yet, somehow, a couple of generations later we're ready to do without logic. Except, that's what our computers run on. But who knows, maybe quantum computers will run not on logic circuits but on hardware neural net implementations.
It will take petabytes to train a "helo world" program and you'll have to wait for a couple of hours for it to learn, but it will be worth the pain, just to get rid of that blasted old logic that "doesn't scale".
The logical part is our interpretation of phenomena not necessarily "out there".
We are pattern recognizing feedback loops the rest is interpretation IMO.
No one is arguing whether we should or shouldn't be using programming languages or hardware based in logic.
Now-
>> this information has to be manually curated (which doesn't scale)
Aye, that's exactly right. I know this first-hand. I study machine learning, so I'm not really-really against it (it works, so). My hunch though is that although we've made a lot of progress with what Hinton calls "analogy" (and what I think he means to call "approximation") we haven't gone _that_ far and every time we made some progress we've eventually hit a wall.
Take language for instance: first there was the Bag of Words model, language as independent words. That worked for a while, but eventually it hit a wall. Now we got word embeddings. It gives great results and yes, you can do arithmetic with words, which is major, but then we still haven't made real progress.
And note that the top-rated NLP system currently, Watson (Schoher's grad student obviously didn't make something as performant as that) combines statistical learning with knowledge engineering. They use a statistical parser to fill in frame [1] slots. Each approach handles different aspects of the problem. Both together cover more of the solution than each on its own can ever hope to. If you want to do really well, you need to pick the best of both worlds, approximation and knowledge, together.
[1] https://en.wikipedia.org/wiki/Frame_%28artificial_intelligen...
I was reading Lin & Wu [1] today; have a look in there for this bit:
to create 3000 clusters among 20 million phrases using 3-word windows, each K-Means iteration takes about 20 minutes on 1000 CPUs. Without using the indexing technique in Section 2.4, each iteration takes about 4 times as long.
So, sure, clustering scales... if you have a um, cluster, of 1000 CPUs lying around to spread it over.
You're saying something fundamentally wrong here. There are classically efficient algorithms (linear and sub-linear in time) for first order logic with search, as in unification, RETE etc, whereas all the machine learning algorithms are typically starting at quadratic times and then increase in crazy, mad, insane rates of complexity that would never be accepted in any other field.
It's machine learning that "doesn't scale", not logic.
The brain is already wired to have certain types of perceptions, so by the time a baby is born that child is just looking for that entity that fits the mommy definition, not trying to define a that it will have a caregiver and the optimal strategy is to bond to that caregiver.
Anyway, there are lots of examples where what you are saying holds, kids drawing general conclusions from limited data, but we also have to remember that the human brain is not a blank slate, it is already highly specialized at the time of birth.
A structure that evolved over time, amusingly, essentially by exactly the kind of "trying stuff at random" algorithm the grandparent is complaining about.
Well ish. Machine learning algorithms are cutting it very close to brute force most of the time.
Edit: My point is not about imprinting (which I don't claim is restricted to humans). I'm claiming that what we consider "human-specific" occupies remarkably little genetic memory. To understand how human language acquisition works with so little genetic context, we need to study humans and not NLP algorithms.
The article was absolutely not saying that we should abandon logic. It simply said that our efforts in building any sort of "understanding" with computers have been more fruitful when relying on analogy and classification as a basis rather than "pure" logic.
Making a distinction between "reasoning by analogy" and "reasoning with logic" is actually pretty weird. Analogy is also (informal) logic - drawing attention to the common characteristics of objects and reasoning about their therefore shared attributes. What Hinton really means, at least as I understand it, is "reasoning with approximation" not by analogy.
But, approximation is what you do when you _can't_ do logic. It's your recourse when you can't do the optimal, which is indeed, reasoning with logic according to knowledge.
Hinton is advocating that, because we managed to get ahead with some hard problems by approximating solutions, we should forget about ever getting a good understanding of the problems. But that's dangerous, defeatist and goes against the way we've progressed pretty much throughout all of human civilisation so far.
The statement was descriptive, not prescriptive. The claim is that the core of our thinking is analogy-based. It doesn't say we are incapable of logical reasoning, or that it is undesirable.
If one agrees with that descriptive statement of human reasoning, then to make an artificial system that does similar things one should have it also reason by analogy rather than logically at its base.
It's a statement of opinion that declares a strongly held view with the clear intent to influence the views of others.
Frex, when theists say that God forbids premarital sex, they don't only say it because they believe it to be true, they say it because they wish you to adhere to their own beliefs.
One thing that I think will be challenging is that language has observer depending meaning. The same statement might have a completely different meaning to someone with a different experience, or made in a different context, or stated by a different person. Games like Chess and Go have observer independent solutions. The winner is the same no matter who/what plays the game.
Determining the meaning of a sentence is a problem where the real answer depends on observer dependent perspective and therefore will need a completely different way to measure success compared to more 'mathematical' tasks like Go. Trying to program a machine to account for this kind of personal experience that humans have, as well as for individual differences between people will be quite challenging I think. I also think that the most significant advances will come from cross cutting academic disciplines like Psychology, Linguistics, and Philosophy of Language.
As we saw with the recent AlphaGo matches, not understanding your opponent (always assuming you're playing against a clone of yourself, as AlphaGo does), can cause objectively suboptimal play. You just have no concept of a "trick play" -- in your world, everything you see, the opponent can see too!
You could argue that with a little bit better theory of mind [1], AlphaGo could have executed a better strategy for its comeback. Instead of playing out obviously broken ladders, hoping the opponent will make a trivial mistake (ha), it was objectively strong enough to devise a cleverer plan, playing to the human's actual weaknesses. It didn't, and lost.
So what you write is true, and only more true in the imperfect, fuzzy and intrinsically deceitful world of NLP.
The only trick play you can do in a game of perfect information is exactly to hope that your opponent will make a trivial mistake in response to your move.
For wodenokoto: you can read more complexity and the difference between "theoretical perfection" vs "practical execution" here [1] (especially the Scott Aaronson’s essay linked there).
1) it's not complex at all.
2) same as 1, with "in 99.99...% of the cases" appended to it.
Most of the problems in AI are a scaling problem. The reason you're not yet seeing androids taking over the economy are twofold:
Portable energy. Our best technologies are not the equal of the human body when it comes to how much power can be stored. But they are constantly improving, even if we'll need another tenfold increase in energy density (less if we make mobile robots gasoline powered, which is impractical for other reasons, more if do the battery + electrical motors thing everyone wants). This means humans are cheaper and easier for a lot of tasks.
Control. The human body (not including face and face-adjecent muscles) has about 300 actuators. That means controlling a human body means controlling 300 individual motors at the same time in a useful way (and every motor affects every other motor. Moving your hand forward means adjusting the power your little toe is applying to the ground to maintain balance). The state of the art is maybe 10-dimensional control (10 interdependent actuators), which can be increased to maybe 20 if the problem can be split into subproblems (e.g. cooperating robots, or parts of the robot that are attached, so imbalance cannot occur). In some ways every extra dimension adds an order of magnitude to the complexity, so it's not like we'll get there in 10 years. But in a century ... probably.
The thing about these problems is that they are problems of degree. Just like today's AI algorithms were known in the 1960's. So why didn't we have live speech transcription in the 1960s ? We all know the answer : processing power was 50 orders of magnitude less than we needed. It was a problem of degree. We knew the problem, and had at least a good guess at what to do, but couldn't contemplate that actually acting on those ideas would make any serious progress due to the limitations of processing speeds at those times. If you wanted a 1960s computer to do a 50x50 matrix multiplication, you'd be waiting days. Doing millions of them was therefore considered useless.
Even this is being gentle. A good case can be made that every component of these algorithms was known once differentiation was formalized, but nobody put them together, not because they didn't realize it could work, but because it was useless : doing things this way would have been incredibly inefficient compared to the then normal ways of doing things. E.g. finding formulas and constants by small adjustments on large chains of partial derivatives (ie. "Deep learning") is something that Isaac Newton knew how to do. It's just he would declare you totally mad for doing that. One might criticize that there were a few holes in the mathematical understanding of matrices in Newton's day, I wouldn't doubt that had he had a reason to really look into those problems, he would have fixed them. But he was more than 50 orders of magnitudes of : doing a single 50x50 matrix multiplication with any amount of resources then available wouldn't have finished before he died.
Now machine learning tutorials tell you to run unrolled 30x30 LSTM expansions on audio data and unroll for 50 datapoints. And then doing that millions of times (tens of thousands of times on a few thousand samples). You can expect this to run on the computer you're reading this on in a matter of hours.
But they were not. To take recurrent networks as an example, active work on RNNs only started in the 80ies. LSTMs, which are one of the solutions to the vanishing gradient problem were only discovered in 1997 by Sepp Hochreiter.
Of course, computational power is one of the limiting factors, but saying that it is the only thing that held AI back is disingenuous and dismisses the hard work of a couple of generations in AI research.
Edit: and Schmidhuber
Aye. Hochreiter and the One Who Shall Not be Named. Let's not forget him again, eh?
Does this mean anything more substantial than 'they knew about neural networks back then'.
> Now machine learning tutorials tell you to run unrolled 30x30 LSTM expansions
Like the LSTM algorithm that wasn't published until 1997?
> So why didn't we have live speech transcription in the 1960s ? We all know the answer : processing power was 50 orders of magnitude less than we needed.
You are seriously underestimating the challenges inherent in the physical sensors required to perform things like speech recognition. Having done some AI/ML on surveillance radar I can tell you that getting reliable data on the physical phenomena is a big part of the challenge in teaching machines to learn and recognize what is going on.
Maybe it means we need to hook up our NLP pipeline to deep-visual networks and knowledge-graphs in order to ask "Did the message of that sentence make sense, and if not, should check a similarly sounding sentence or a different parsing of the sentence in order to make it more cohesive"
I do think we have NLP algorithms that can handle garden paths without understanding the sentence.
A state of the art Japanese POS-tagger for example has a dictionary of all previously seen word/pos pairs and their transition probabilities. Sine there are no spaces in Japanese, the possible combinations of word/pos pairs you can lay out that matches the input sentence is quite big, and gives you a fairly large graph that you need to find the globally best path through.
This means the algorithm does walk down the garden path, but at some point decides that there is a better path that will get you to the end of the sentence.
If you go to http://www.atilika.org, scroll down a bit, select the "viterbi" radio button, paste 私は日本人です into the text box and click tokenise, you can see an example of such a graph.
I kind of see statistical parsing as another "parlor trick" type of AI. Impressive and hard to achieve but somehow separate from true intelligence. It's hard to formally describe this distinction which is why I think that advances will come from cross cutting academic departments and reaching out from the "throw more math and more processors at the problem approach"
To achieve a true AI we will have to advance not just our understanding of algorithms, but our understanding of what it means to be intelligent which is a concept that I think most people lack any coherent definition of.
The idea with word embeddings is that they provide a context for words, which informs their meaning.
It goes back to Zellig Harris and the observation that "constituents of the same type can be replaced by each other" [1], or in other words: words that occur in the same context have the same meaning.
[1] Zellig Harris. 1951. Methods in Structural Linguistics.
https://cs.stanford.edu/~quocle/paragraph_vector.pdf
(This is to say there is such a thing as an unsupervised method for creating features that encode meaning, and it's being used today for various tasks :] )
To me, Google search with NLP is at best barely better than search with some kind of query language (and often worse imo). Those phone-based robobilling system now can mostly understand what you say but are still essentially only a notch above keypad menus.
The thing with complex syntax-based meaning in human language is that really is only worth the trouble if one is setting up an ongoing relationship with an other. A very smart human with vast knowledge still couldn't give great answers to single-sentence questions from people he'd never seen before and who would never see him again.
And ongoing language interactions actually tend to connect multiple aspect of human activity making divisions harder. So metaphorically we may have a few very small fruit and more or less one very large one.
I have worked off-and-on on a project to detect metaphors and identify the concrete meaning the metaphor is attempting to convey. Amongst many other things that project taught me that it is nearly impossible to get more than five linguists to agree on what a metaphor is in the first place.
One project I completed was able to really excel at keyword matching simply by building a huge dictionary of words, in a literal sense: a dictionary of contextually relevant words to a phrase, generated by very large texts.
I think baby steps are the key to getting further with NLP. For reference: http://nlp.stanford.edu/fsnlp/
Yeah, I know the work he's talking about. It's the one related to this dataset:
https://archive.ics.uci.edu/ml/datasets/Kinship
From that page:
Creator:
Geoff Hinton
Donor:
J. Ross Quinlan
Data Set Information:
This relational database consists of 24 unique names in two families (they have equivalent structures). Hinton used one unique output unit for each person and was interested in predicting the following relations: wife, husband, mother, father, daughter, son, sister, brother, aunt, uncle, niece, and nephew. Hinton used 104 input-output vector pairs (from a space of 12x24=288 possible pairs). The prediction task is as follows: given a name and a relation, have the outputs be on for only those individuals (among the 24) that satisfy the relation. The outputs for all other individuals should be off.
Hinton's results: Using 100 vectors as input and 4 for testing, his results on two passes yielded 7 correct responses out of 8. His network of 36 input units, 3 layers of hidden units, and 24 output units used 500 sweeps of the training set during training.
Quinlan's results: Using FOIL, he repeated the experiment 20 times (rather than Hinton's 2 times). FOIL was correct 78 out of 80 times on the test cases.
And yet, if you have a wee look at Hinton's publication on Rexa, there's 43 citations, while there's a single one on Quinlan's (from Muggleton, duh).
So, you know, maybe it's not logic and reasoning that's the problem here, rather a certain tendency to drum up results of neural models even when they don't do any better than other techniques.
But, really, it doesn't matter. Google has the airwaves (so to speak). No matter what happens anywhere else, in academia or business, their stuff is going to be publicised the most and that's what we all have to deal with.
One thing that bothers me about ANN hype is that those models are very resource and data hungry. Which suits Google perfectly, but will keep that kind of AI out of people's hands/computers for decades to come.
father(Christopher, Arthur)
father(Christopher, Victoria)
father(Andrew, James)
father(Andrew, Jennifer)
...
mother(Penelope, Arthur)
mother(Penelope, Victoria)
mother(Christine, James)
...
This is about as close to a "logical input" as one can possibly get. One would naturally expect a logic-based system to handle this very well, so the fact that it indeed does is not that interesting.Moreover, IMHO, it is not clear how we could extend a logic-based system to handle the messiness of natural language. (Well, perhaps it was also unclear how to do that with word vectors, when Hinton wrote that paper, but now it's pretty clear.)
Not saying word vectors are the "ultimate" solution, but logic isn't even at the level of word vectors.
>> it is not clear how we could extend a logic-based system to handle the messiness of natural language.
There's a lot of recent stuff in Muggleton's Latest Advances in Inductive Logic Programming and a ton of work from others, before and after that.
There's no problem in handling "messiness" with logic. Like I note elsewhere, your computer is a logic machine and it handles messiness just fine (no offense meant).
Since the ACM has professional editors, I was surprised that they would twice misrepresent the example "king – man + woman = queen" as "kingman+woman=queen" in the article. At least they spelled "Hello, World" right, even if they couldn't bring themselves to add the "!".
It looks like the problem is that in the PDF version, the phrase happens to be hyphenated at the "minus sign" in both usages (http://delivery.acm.org/10.1145/2880000/2874915/p13-goth.pdf) [1] although one might hope this is something an editor would have checked.
[1] Looks like the ACM wants to you click on the PDF link yourself, from the "View As" bar in the text version.
http://webcache.googleusercontent.com/search?q=cache:145V9qm...
Well, for some that's certainly the case.