There may be a way to build it into the loss function so that it happens before the quantization, right?
(and holy crap, look how fast the HN conservatives are getting to my comment)
(and holy crap, look how fast the HN conservatives are getting to my comment)
But any interesting release of NLP data has the potential to affect the way the field progresses, so take it as a compliment that I consider this an interesting release of NLP data. That's why I'm asking you to actively consider the downstream effects of word vectors and find out if you can make them better.