The original paper didn't receive the attention I thought it would, but I continue to think this is a fascinating result which has deep implications for machine learning and for linguistics.
The original paper didn't receive the attention I thought it would, but I continue to think this is a fascinating result which has deep implications for machine learning and for linguistics.
It might seem obvious that things should be done that way, but if you consider servers hosting lots of different websites on the same physical machine, or data structures spread out over several memory allocations held together by pointers, it's clear that there are other possibilities. So it does seem to be specific to the way humans use language.
And because human language has this property of co-occurrences corresponding to relatedness in meaning, you can represent the meaning of a word by building a model that only predicts the probability that two words occur together.
There has been much in the way of discussion for, and concern about (for example, the Deconstructionist movement), these ideas, for the past 60 years or so. And a bit of practical exploration in the field of child development.
Interestingly, the fact that embeddings between languages seem to share some common shapes (per the linked paper in this thread), would seem to suggest that A) Fundamentally, most languages have the same deep structure, whether through coincidence or common evolutionary root. or B) The brain has a hard-wired structure for language, evolved alongside the development of language itself. The Chomskian Language Acquisition Device. We're not born tabula rasa, we've got some hardcoding indicating how we're going to understand things
Or a little of A, a little of B maybe, as it does end up being a boostrapping problem.
Normally this technique wouldn't be useful, because it's overfitting a specific training set. (If you make space X as similar as possible to space Y, then this mapping from X to Y is only useful for X to Y – it can't generalize to other situations, which is often the goal of an ML model.)
But since the task is "Translate from English to Italian," and since all languages have similar embedding structures (Zipf mystery), overfitting is exactly what we want: we want, for any given English phrase, to find the closest-fitting mapping to a corresponding Italian phrase.
The more I learn about ML and data mining, the more I'm astounded by how clever many of the techniques are, and how much artistry is involved. You have to be clever to make a certain model perform well in a certain domain. If you want to make a stock trading bot, you can't randomly subdivide stock market data into e.g. 70% training data and 30% test data, because the data is ordered by time. You have to use the past 3 years of stock market data as a training set, and validate it against the subsequent 1 year of market data.
I really like ML because the techniques applicable for training a stock market bot seem unrelated to the algorithms for doing unsupervised machine translation, which differ from how to model credit fraud, which are no doubt different from how to build a dota 2 bot. :)