Show HN: Language games – simple games made with word vectors
github.com
github.com
Same with any other language models and NLP related stuff. It's hard to find anything valuable.
I don't think there's been much happening around word vectors specifically, but it's still worth knowing about.
Actually facebook research are pretty good at gathering things, their ParlAI project also gathers together other sources for conversational stuff.
Furthermore, given a copy of the text from wikipedia, it only takes 2 days of computation on a laptop with a modern GPU to regenerate them from scratch for any language having a wikipedia (I did that using two R packages a few years ago, I expect that it would be even simpler nowadays).
(And it's "pedant", not "pendant").
I did something similar when I was still at school. Scripts which emulated Codenames [1], and providing hints for given set of words [2].
I believe that is possible extension to your project - emulate Codenames games:)
I’ve also played around with word vector games — Robot Mind Meld [1] has you and a robot working together to converge on the same word.
Player 0 it is your turn!
The word you are trying to match in meaning is:
antepenultimate
Your char list is:
['i', 'g', 'j', 'u', 'b']
Please input a valid word made from some or all of the provided characters
What's the right answer?It calculates the distance in vector space from your word to the given word.
I think it helps if you understand what Word2Vec is, as then it becomes clear what is going on.
Is it just "words that vaguely, extremely distantly, have some relation, as determined by the neural net"? And your job is to work out which extremely-distant word is slightly closer? (I noticed "gubi" scored very, very slightly more points than "bug" in the demo.
Essentially, all words are embedded in an N-dimentional space. Then you can simply measure the distance between words in the word-cloud. The exact method of how it decided what coordinates to assign to each word in the cloud isn't important to the game.
The fastText vectors (see below) also use subword embeddings so I think they had potentially better results. I used FT for some Chinese stuff and I think it worked better for that since chinese characters are so much more important than latin scripts.
Seems magnitude python API also has some other features like POS tagging, I wonder how that compares to say spaCy.