Learning Word Vectors for 157 Languages
arxiv.org
arxiv.org
Word vectors are fascinating representations. There is a huge amount of information and nuance captured in them. You can use them directly for topic retrieval (using annoy or another optimised vector index), or feed them into a classifier such as those in the Sklearn library. All types of neural nets: fully connected, recurrent and convolutional can be applied on word vectors.
Pre-trained data is valuable, and you really aren't going to re-learn it all from your data, so why throw it out?
This is covered in fast.ai 's new course 1.
Regarding the punctuation, as pointed out in another comment, these tokens might also be useful for some applications (and they are easy to filter out if you don't need them).
And yes I realize this is a really odd question :)
My first guess would be emojis ;)
For example, here there are all kinds of useful things we can do with these 157 sets of word vectors, but Cantonese escaped the list because most of its transactions happen off the page.
"One of the challenges for Natural Language Processing (NLP) systems is the question of how to represent input such that the network runs quickly, but also learns well. It's possible to represent each word as a one hot vector, but that's computationally slow. It's also possible to represent each word as some number, but then lots of words look very similar.
Instead, why not use a mix? Introducing word embeddings. We'll represent each word as a n-dimensional vector, with each dimension representing a trait about the word. For example, "fruit" might be represented as {food: 0.99, gender: -0.05, size: 0.2}, and "king" might be represented as {food: -0.9, gender: 0.92, size: 0.56}." [Quoted from MuffinTech.org] [See v1n337 for caveats. [0]]
Two similar words should have similar word vectors, like "apple" and "peach". If we learn some fact about apples, like "Humans eat apples.", then we can easily generalize that to peaches, pineapples, etc...
Let's tie this back to the research. Since we have word vectors for many languages now, that makes it easier for us to build NLP systems in other languages. For example, if we wanted to build an English->French translator.
Although, I suppose if we treat "apple" and "Apple" as different words, that would help.
Fun fact: One of the current NLP problems is detecting which words are names. Apparently it's really tough, especially with Twitter data!
Unfortunately I see a problem in having to specify an exact position per word. If you think of the position of english "Apple" in the Spanish word space as a distribution instead of a specific location, then it ideally should be a two-mode distribution, with one peak next to Apple and one peak next to manzana. If you must use a normal distribution, the variance must be wide enough to cover both words -- a huge problem, since (a) that assigns a lot of probable values to one word and (b) the mean value (expected value) lies between them, not at the semantic location of "apple" at all.
An interesting paper looked at how these associations changed over time [1]. It was also featured recently on The Morning Paper [2], in case you prefer a summary with added context.
Although those ambiguities make things a bit more difficult, you can usually leave the job of disentangling them to a later stage in the language-modeling process, which will have more context it can use to disambiguate which word sense was used.
[1] https://arxiv.org/abs/1703.00607
[2] https://blog.acolyer.org/2018/02/22/dynamic-word-embeddings-...
EDIT: Indeed, this is old data from a previous publication. It appears they have not actually made the new data public yet.
If you are interested in more, check out these excellent reviews by Adrian Colyer posted in The Morning Paper.
https://blog.acolyer.org/2016/04/21/the-amazing-power-of-wor...
word vectors are vector representations of each word in the vocabulary. Here they are learned by a neural net. the length of the vector is the # of features. Just for intuition, one feature of a word the NN could learn is the gender of a word, and so on.
But the features aren't individually interpretable, in practice. For instance, the 'gender' of a word may have it's signal scattered over several features/dimensions of the learned vector.
"fruit": {food: 0.99, gender: -0.05, size: 0.2}
"king": {food: -0.9, gender: 0.92, size: 0.56}
Building off of what v1n337 stated, though, axis can easily be skewed and rotated such that they're still interpretable, just not obviously so.
This has very interesting possibilities: you can complete sentences where a word is missing (you have a context, so you can search for the best matching word vector), use it in text autocorrection tasks, and other classic natural language processing problems.
Learned word representations are also coherent between them, so you can use them to make analogies (the distance from 'Spain' to 'Madrid' is similar to the distance from 'France' to 'Paris'), so they implicitly hold some of the semantic info between words.
It can also be used to find related words. Synonyms, antonyms and related words have similar representations. For example, 'facebook', 'twitter' and 'instagram' have similar vectors (vectors with similar directions). But you can also try with famous musician or band names, tech related terms, etc.
Finally, word vectors, unlike other language models, can store representations for large windows and vocabulary sizes in a few GB, which is another useful property in certain situations, and makes them easy to handle.