Would someone please explain in simple language what this is, and why it's cool?
Would someone please explain in simple language what this is, and why it's cool?
Word2Vec is one of the algorithms to do this. Given a bunch of text (like Wikipedia) it turns words into vectors. These vectors have interesting properties like:
vector("man") - vector("woman") + vector("queen") = vector("king")
and
distance(vector("man"), vector("woman)) < distance(vector("man"), vector("cat"))
What Word2Bits does is make sure that the numbers that represent a word is limited to just 2 values (-.333 and +.333). This reduces the amount of storage the vectors take and surprisingly improves accuracy in some scenarios.
If you're interested in learning more, check out http://colah.github.io/posts/2014-07-NLP-RNNs-Representation... which has a lot more details about representations in deep learning!
A big problem with NLP is understanding the semantic associations between words (or even lemmas. Lemmas in this context refer to different meanings of the same word, like a baseball bat vs. a vampire bat). For example "run" and "sprint" are similar in meaning but convey different connotations; kings and queens are both high-level monarchs but we need to encode the difference in gender between them for a true semantic understanding. The problem is that the information in words themselves don't accurately convey all of this information through their string (or spoken) representations. Even dictionary definitions lack explanations of connotations or subtle differences in context, and furthermore aren't easily explained to a computer
Word2vec is an algorithm that maps words to vectors, which can be compared to one another to analyze semantic meaning. These are typically high-dimensional (e.g. several hundred dimensions). When words are used often together, or in similar contexts, they are embedded within the vector space closer to each other. The idea is that words like "computer" and "digital" are placed closer together than "inference" and "paddling".
Usually these vector mappings are represented as something like a python dictionary with each key in the dictionary corresponding to some token or word appearing at least one (maybe more) times in some set of training data. As you can imagine these can be quite large if the vocabulary of the training data is diverse, and due to being in such a high-dimensional space, the precision of a vector entry may not need to be as high as doubles or floats encode. Floats are 32 bits. The authors of this paper/repo figured out a way to quantize vector entries into representations with smaller numbers of bits, which can be used in storage to make saved word2vec models even smaller. This is really useful because a big problem with running word2vec instances is that they can take up space on the order of gigabytes. I haven't read it all yet but it seems the big innovation might have been figuring out a way to work with these quantized word vectors in-memory without losing much performance
Edit: seems it may not work in-memory
But when you save them to disk every value is either -1/3 or +1/3 so one could encode the word vectors in binary. This can lead to reducing memory usage during application time if you kept the word vectors in this compressed format (though you'd need to write a decode function in tensorflow or pytorch to take a sequence of bits corresponding to a word and convert it into a vector of -1/3s and +1/3s)
Intuitively I can suggest the following: One of the problems is that even relatively huge corpuses like wikipedia can't deal with high dimensionality that well because even they likely lack the sheer size and diversity of content to "flesh it out", so to speak. My intuition tells me that this is due to the training process overfitting "clusters" of information together due to the size of the model. If you have too many dimensions, there's a lot of space for clusters to form, and with that will come some loss of semantic differentiability between clusters - or at least my hunch tells me so. You definitely want "computer" to be more associated with "mail" than "taupe" but if the frequencies of their associations are small they'll essentially be interpreted as noise or overfit. One thing to note is that word2vec embeddings are trained using a shallow neural net, and it's entirely possible for it to be a generic ML problem of too-many-parameters/bad network topology given the input when dimensions get too high.
With dimensions in this context it can be easy to forget that each additional dimension added can (potentially) add an order of complexity to the model - a smaller model lies on a hyperplane in the new vector space. What may happen is that an added dimension (by added, I mean before training, not after) might add some beneficial complexity for a specific subset of the model; e.g. if we were originally in very low dimensions, adding one dimension may allow king/queen and boy/girl to separate in the vector space based on gender( although in reality you can't create correspondences to individual dimensions and properties like this usually) but simply lead to noise or overfitting in other subsets. I think in very high dimensions this overfitting is likely to manifest itself in either too much clustering or strangely high similarity words resulting from outliers/noise in the data.
I've never really seen thousands of dimensions used in the wild, but I don't know of any papers that explicitly compare performance among different dimensions (word2vec is hard to evaluate with a single metric, though). Perhaps once you get to internal Google or Facebook levels of big data you could use distributed computing to make it work, but again, I haven't seen references to that.
In geometric terms, think of a circle embedded in a square. As the number of dimensions increases (e.g. a sphere embedded in a cube, a 4-dimensional ball embedded in a 4-dimensional cube, etc), most of the volume in the cube is outside the radius of the sphere. Any vectors you have effectively become sparse.
Basically this means that, while you need large dimensionality to model complex relations, high dimensionality makes it difficult to model these relationships (and need exponentially more data).
The wiki page on Word2vec seems helpful.