Unsupervised machine translation
code.fb.com
code.fb.com
The original paper didn't receive the attention I thought it would, but I continue to think this is a fascinating result which has deep implications for machine learning and for linguistics.
It might seem obvious that things should be done that way, but if you consider servers hosting lots of different websites on the same physical machine, or data structures spread out over several memory allocations held together by pointers, it's clear that there are other possibilities. So it does seem to be specific to the way humans use language.
And because human language has this property of co-occurrences corresponding to relatedness in meaning, you can represent the meaning of a word by building a model that only predicts the probability that two words occur together.
There has been much in the way of discussion for, and concern about (for example, the Deconstructionist movement), these ideas, for the past 60 years or so. And a bit of practical exploration in the field of child development.
Interestingly, the fact that embeddings between languages seem to share some common shapes (per the linked paper in this thread), would seem to suggest that A) Fundamentally, most languages have the same deep structure, whether through coincidence or common evolutionary root. or B) The brain has a hard-wired structure for language, evolved alongside the development of language itself. The Chomskian Language Acquisition Device. We're not born tabula rasa, we've got some hardcoding indicating how we're going to understand things
Or a little of A, a little of B maybe, as it does end up being a boostrapping problem.
Normally this technique wouldn't be useful, because it's overfitting a specific training set. (If you make space X as similar as possible to space Y, then this mapping from X to Y is only useful for X to Y – it can't generalize to other situations, which is often the goal of an ML model.)
But since the task is "Translate from English to Italian," and since all languages have similar embedding structures (Zipf mystery), overfitting is exactly what we want: we want, for any given English phrase, to find the closest-fitting mapping to a corresponding Italian phrase.
The more I learn about ML and data mining, the more I'm astounded by how clever many of the techniques are, and how much artistry is involved. You have to be clever to make a certain model perform well in a certain domain. If you want to make a stock trading bot, you can't randomly subdivide stock market data into e.g. 70% training data and 30% test data, because the data is ordered by time. You have to use the past 3 years of stock market data as a training set, and validate it against the subsequent 1 year of market data.
I really like ML because the techniques applicable for training a stock market bot seem unrelated to the algorithms for doing unsupervised machine translation, which differ from how to model credit fraud, which are no doubt different from how to build a dota 2 bot. :)
I'm not sure I follow the qualm you are trying to get across. Are you saying you disagree with the term 'unsupervised' because unsupervised algorithms still bake in human assumptions (like a human-designed word embedding model) so that's essentially still supervision?
The obsession with 'unsupervised' learning as the quintessential technique is about getting better results for less money/effort. The premise is that we assume deep models tend to scale up in accuracy as training data size increases, so we always want larger datasets to increase accuracy. But creating labeled data takes a linear amount of human effort ($$$) as the dataset size increases. At a certain point, creating more labeled data to improve a model is not cost effective or maybe even impossible.
Unlabeled data can be acquired nearly for free in nearly unlimited quantities in many cases. So if we can use unlabeled data instead (even if it requires complex pre-processing like CBOW embedding models which essentially turn bits of the unlabeled data into it's own label) your final results-per-dollar-invested goes through the roof over supervised learning. That's the obsession. It's not about literally no supervision being involved in the process. It's about driving down the cost of data acquisition while driving up the percent of the world's available data you can use for training a model.
I apologize in advance if I'm missing the point you are making.
Where I disagree with you is that the obsession is purely driven from "results-per-dollar-invested", at least in the academic world. That being said, unsupervised learning is a great tool and definitely worthy of research.
To summarize, my comment was completely tangential of this paper (the authors make no such claims). It was more of a stream of consciousness comment that arose because I envisioned someone reading the paper and saying "see!, unsupervised learning leads to real understanding, no humans needed!"
In one of my previous companies we used this technique to hire pair of translators: we gave translator pairs such task, and hired pair which reconstructed original text more closely.
2. rotate embeddings space of two languages for optimal alignment, assuming frequency and neighborhoods of word embeddings are more-less the same in any language
3. iteratively minimize difference in bidirectional translations
Does anyone knows some references / corpus in English language for this?
To translate whale song, you need data that corresponds to what the whales are singing about.
Wouldn't the method work better with n-gram embeddings, where n=3 or 4?
And since you are simply learning a rotation matrix, there is no risk of overfitting.