Why Deep Learning Cannot Be Applied to Natural Languages Easily
linkedin.com
linkedin.com
It's true that neural networks can not be easily applied to natural languages, but there are less obvious ways of applying them (namely Embeddings, LSTMs, Attention) and provided you have enough data and computational resources they give so much better results than any other method that it doesn't even help anymore to combine them.
[1] https://arxiv.org/abs/1609.08144
[2] https://arxiv.org/abs/1611.04558
EDIT: I also don't want to give the impression that Deep Learning can solve every NLP problem. We are still far away from passing the Turing test. It's true as well in my opinion that Google's Machine Translation is oversold. It's best at what it's trained and evaluated for: translating individual sentences from news sources.
EDIT 2: There are some tasks where traditional methods can work better, e.g. text classification on long documents. That's mainly because deep learning methods are too expensive computationally.
Honestly that seems to be the default for technical posts on linkedin.
In general, RNNs are especially fit for discrete sequences. And their continuous representation is actually an advantage (so they can see that two words are similar or analogous). BTW: see my draft on word2vec: http://p.migdal.pl/2016/12/30/why-do-word2vec-analogies-work...
He says: "Why Deep Learning cannot be Applied to Natural Languages Easily"
But I say: "It's true that neural networks can not be easily applied to natural languages," Take that!
But at the same time they say "It's true as well in my opinion that Google's Machine Translation is oversold" so the "they don't know what they're talking about except for all the points that I agree with them on" tone seems a bit silly.
"A Neural Network for Machine Translation, at Production Scale"[1] says:
Machine translation is by no means solved. GNMT can still make significant errors that a human translator would never make, like dropping words and mistranslating proper names or rare terms, and translating sentences in isolation rather than considering the context of the paragraph or page. There is still a lot of work we can do to serve our users better. However, GNMT represents a significant milestone. We would like to celebrate it with the many researchers and engineers—both within Google and the wider community—who have contributed to this direction of research in the past few years.
That seems fair to me.
"Found in translation: More accurate, fluent sentences in Google Translate"[2] says:
Today’s step towards Neural Machine Translation is a significant milestone for Google Translate, but there’s always more work to do and we’ll continue to learn over time. We’ll also continue to rely on Translate Community, where language loving multilingual speakers can help share their language by contributing and reviewing translations.
"Zero-Shot Translation with Google’s Multilingual Neural Machine Translation System"[3] says:
In September, we announced that Google Translate is switching to a new system called Google Neural Machine Translation (GNMT), an end-to-end learning framework that learns from millions of examples, and provided significant improvements in translation quality.
[1] https://research.googleblog.com/2016/09/a-neural-network-for...
[2] https://blog.google/products/translate/found-translation-mor...
[3] https://research.googleblog.com/2016/11/zero-shot-translatio...
This isn't cherry picking - you can try it yourself.
Randomly selecting the 2nd paragraph from https://www.hrw.org/fr/news/2017/01/02/israel/palestine-des-...
Human Rights Watch a documenté de nombreuses déclarations faites depuis octobre 2015 par des personnalités politiques israéliennes de haut rang, dont les ministres de la Police et de la Défense, appelant la police et les forces armées à tirer sur des individus suspectés d'attentats pour les tuer, avant même de déterminer si le recours à la force létale est, ou non, absolument nécessaire pour protéger des vies.
Choose which is automatically translated:
Human Rights Watch has documented numerous statements since October 2015, by senior Israeli politicians, including the police minister and defense minister, calling on police and soldiers to shoot to kill suspected attackers, irrespective of whether lethal force is actually strictly necessary to protect life.
vs
Human Rights Watch has documented numerous statements made since October 2015 by high-ranking Israeli politicians, including police and defense ministers, calling on police and armed forces to shoot at suspected terrorists even before deciding whether or not the use of lethal force is absolutely necessary to protect lives.
One of these in manually translated by an expert, one is by Google Translate (clicking on the "English" link on the page linked above will show the expert translation).
You can find other document collections to test here: http://www.miis.edu/academics/library/find/guides/translatin...
To honest it works better than I thought for French > English, but I when I tried some random French websites (not news) the results were not as good as this example and nowhere near human translations.
No, the "perfect score" is an "averaged" translation by multiple humans, measured by BLEU score.
The premise of the article is valid though, in that NLP is a hard problem. The reason is partly because NLP is ill-defined; how do you define language understanding?
NNs are very effective at learning mappings of y=f(X), given enough examples. One of the reasons that they're so effective at modelling speech, vision, translation, etc., is that such mappings exist in high volumes. Because of the above-mentioned ambiguity of NLP, it's harder to come up with such pairs for 'understanding' a language. How do you come up with a dataset of sentences and their 'meaning'? Probably the best you could do is to map them to some action. And critics will readily disregard such attempts as 'not really NLP'.
And there have been attempts to ascribe a semantics to natural language from text (for ex. see CCG grammars). The datasets are not as big as for vision tho, yes. But I'm not convinced that we need such explicit datasets to be able solve this problem.
If he's going to make technical arguments, then they should at the very least be relevant and correct.
There are just a lot of things that have to be figured out still.
+ Different time scales. There is semantics on a sentence level while there is also semantics on a plot level. It's convenient to know key elements from the start of a story if you want to understand the plot. LSTMs are a perfect starting point.
+ When to stop learning. The so-called stability-plasticity dilemma. Our ability to pay attention to what matters might be tightly linked to our capability to forget vast bodies of texts that we just read. Current NNs do not seem to forget correctly. This was the rationale behind ART and ARTMAP (Grossberg) and might enter AI mainstream again soon.
+ Grammar constructions. Some aspects of grammar seem simpler than computer vision, where we also have a lot of structure in the environment, models like things that can be inside of other things, be balanced on top of other things, temporarily occluded by other things, etc. Other aspects seem more complicated, like the pleasantness of a poem. My gut feeling is that some of this gets spilled over from (a) structure in other modalities and (b) idiosyncrasies from our generative system (vocal cords, etc.). In other words, our grammatical preferences might be sampled not only from listening and reading.
+ Emphasis.
Just a few things that might lead to interesting NNs. Contrary to the author I think they are definitely in line with current research.
By question answering based on text. There are many papers doing that.
Deep learning is part of enormous advances in NLP[0], just as it has set records to accuracy in almost every field of machine perception, from vision to audio.
Neural word embeddings like those produced by word2vec[1] make for very useful feature vectors when fed into other neural nets.
The headline of this post should be that NLP is harder than, say, image processing. In fact, for non-specialists, none of it is easy, because tuning hyperparameters is hard.
The kind of NLP that tries to reproduce human-level sentences and understanding is simply a more complex problem, given the plasticity of language.
[0] https://arxiv.org/abs/1611.04558 [1] https://deeplearning4j.org/word2vec
> What is a continuous function (or continuous data)? It is a sequence where each item is related to the one before and one after determined by a process.
So, to paraphrase the article: My reply should actually stop here with one sentence.
[1] Google Translate on the German Wikipedia entry for Weihnachten (X-mas): https://translate.googleusercontent.com/translate_c?depth=1&...
I don't think anyone deny that 'true' translation would require some kind of general intelligence that somehow understands what is being translated, but it seems to be the case that a 'dumb' translation works well enough for a great many use cases, regardless.
He's really just making the Chinese room argument. We have a computer shuffling symbols around according to some rule set, that doesn't know what they mean. I don't think it really matters, though, if it produces a reasonably accurate translation.
I think the counter to Searle's argument isn't really that it doesn't matter as long as the result is close enough. The counter to that is that we don't understand how human intelligence works either. Searle is simply assuming that it's "magic" (or less condescendingly, some sort of metaphysical process) that can't be simulated by algorithmic machine. I think it's far more likely that intelligence is physical and we just don't understand the machinery than it is that it's mystical and cannot in principle ever be understood.
For this article, all that is seemingly unnecessary. He's just saying they won't work well enough to even fake it convincingly. Which is very nearly falsifiable just by running today's algorithms.
But if you're at 75% per cent and want to get to 100%, you need to understand what the problem is with the rest. And if it's 75% of "perfect translation of single written sentences from newspapers or technical litterature", how far along is that towards something "being part of an everyday conversation"?
I studied linguistics at a university where focus was very much spoken language, sociolinguistics, language in context, before moving into (or through) NLP, and the distance between what a statistical machine translation system is able to handle and the stuff I used to work with is very large.
A lot of NLP work now seems to focus on algorithms, but intuitively it seems to me that a much larger issue is the quality of the data, in the sense that humans don't learn language from piles of isolated text and somehow we're expecting machines to do it.. Rext is a lossy encoding of spoken language, even if you try your best to mimic it, but more seriously it does not include the physical context that children encounter language in. The learning situations aren't the same, I don't know why we're expecting the results to be.
What machine translation isn't going to do is capture emotion or style or understand what someone is saying without them really saying it and so on. I would be very surprised if a machine translated novel ever hits the best seller charts. But I bet translators are going to be working from machine translated glosses if they aren't already.
The argument goes - "when you look inside, it's just things pushing at each other. How could that produce perception and the conscious mind?"
So, a failure of imagination and incredulity based on how they understand the world and the mind makes them reject AI. They feel that the special place of the soul was traded for "information processing" which is dry and mechanical - a form of dualism creeping up in our day and age.
I would have felt the same if I didn't learn and use neural networks such as CNNs, RNNs and MLPs. Now I know how simple mechanical systems can recognize patterns and process information to generate complex behavior and I don't feel that "explanatory gap" any more.
Reinforcement learning is a good base for consciousness research - much more precise and with scientific results, not just p-zombies and bat based armchair experimentation. There's a limit where you can go with just pure thinking and then you need to start direct implementation.
That cognitive inference process is what we've formalised as probability theory.
Whenever you do /anything/ your brain may be selecting from a probability distribution over things that can be done immediately.
As for text as continuous data just chuck it in glove, word2vec, lexvec or fasttext. Given enough training you could model the velocity of concepts as they're being introduced to the dataset / model.
Also, on the whole, shallowish learning can be applied to Natural Languages pretty easily. Keras includes a memory network (LSTM with an autoencoder) that averages around 98% accuracy on the bAbI 10k Q/A task.
There is no concrete evidence that the brain does math. If the brain did select things from a probability distribution, then why isn't everyone a math genius.
No one really knows how it works.
As for your second point, assuming that humans are Bayesian, there are many reasons why people would have variability in their mathematical ability, including different priors and differences in the ability to estimate posteriors.
[0] https://scholar.google.com/scholar?q=brain+bayesian&hl=en&bt...
There are people who never manage to learn to do basic elementary-school math with fractions and percentages and yet manage to bet on sports and balance their checkbook because they're unable to translate their intuitive mathematical instincts into abstract formal math.
The brain absolutely does tons of math. It's just well below the level of consciousness. Most people can learn to do analog math problems well within a few percent accuracy. It's symbolic math that is foreign to us.
If the apple falls predictably from the tree, where is the math genius who plans its trajectory? Does the apple itself know physics?
> On the other hand, grammar rules and ontological semantics mastered by the human brain can handle the entire sextillion (since those pages were written by human). If you know how to read and write, the entire sextillion will be understandable to you. This is the horrifying truth between the capabilities of the human brain versus the current state of neural networks.
This bit especially doesn't make any sense. I'm a human who has been reading all my life, does that mean I understand every grammar rule or ontological semantic ever created? Of course not! I barely understand all of them in my own language! My 'neural network' (brain) would need a bit more 'training' (studying) before that could happen. Even more, if all I had read in my life were 10,000 pages (which still may be true).
Last year I tried translating some everyday text from Turkish to English. Complete garbage, you could barely understand even what they were talking about.
I tried it now, albeit with different texts, and there's a world of a difference. Now you can actually understand what the text is saying, even if compared to other languages, I would still classify the Turkish->English translation as awful. Another big difference is that the resulting English text has relatively good grammer, as opposed to the previous version which was a broken English word soup.
I think the author is overlooking the fact that images are not continuous functions but Deep-Learning image-recognition systems have been very successful representing discontinuities in images as hierarchical visual abstractions. In the same way, Deep Learning with recurrent neural networks should be able to learn discontinuities in symbol streams as hierarchical language abstractions, given enough data.