Semantle: guess words based on Word2vec similarity scores
semantle.com
semantle.com
I found it was nearly impossible to enjoy because a single numerical distance metric provides almost no help in figuring out what direction to move to improve your guess, and any change in the distance metric when you're far away from the solution is really just a false signal - i.e. "mother" gives a score of 20, "grandmother" gives a score of "23", but grandmother isn't actually closer to the final word ("malfeasance") in any meaningful way than mother is.
Another issue with it is that the training set comes largely from newspaper articles, which means the semantics it learns sometimes have significant artifacts resulting from associations that are common in news coverage but not the English language in general.
As a result my own experience with it was "think of random guesses endlessly until you end up somewhat close and then be frustrated at what you learn about the embedding when you eventually converge on the answer."
3.8 as "cat" is to "pig"
I think this is largely the point. The game is significantly easier (for me at least) when I limit myself to thinking about what words would have high co-occurrence in English print journalism.
Doesn't this make sense though? A small change like that means you didn't really move, so you didn't get further or nearer the solution. Grandmother and mother are pretty similar.
I have also found it frustrating in semantle when I get something that is "close" but is actually close because it is a direct antonym of the word.
I've had this same idea to use the point cloud visualization to improve my own take on language model games - but I when I experimented with this, I found that even the best 2D representations were still quite bad in the context of trying to help out humans.
I had an idea for a minor improvement where if you guess one word, behind the scenes it would check that word, the plural version, the past tense version, etc. And then show you the highest score from all of those. I find it very frustrating when one word has a certain similarity score but another version of the same word (pluralized, past-tense, etc) has a very different score, whether higher or lower. Some small tweaks like this might make the game much more enjoyable.
Edit: So the way he grabs words for the original is out of the 5000 "most popular" words in English. Still not sure about the Junior version, maybe it is an even more restricted list. Or it looks like it might just be restricted to nouns since it does seem to be easier to guess things close to a physical item.
Mum turned 80 yesterday, sister turned 50 2 weeks ago and I'm 55.75 now.
5 June 2022
We play it like the 20 questions car game from the 1970's, by starting with smalk, mineral, vegetable, but have added a few more class headers, as we call them, chemical, emotion, transport, colour, human, etc....
Junior is usually played the same with easier words, so can be found quicker, Eg I always play animal as my first word and the other day that's what it was so boom! found it in 1 guess!
Hooked!!
Heres the the top list for me: https://imgur.com/AqmQl6A
But also I stopped playing because it wasn't fun anymore.
https://github.com/Hellisotherpeople/Language-games
Also, you can use extensions to word embeddings, namely, sense2vec (https://github.com/explosion/sense2vec), to supercharge games like this!
https://greatfilter.itch.io/lexicode
I am learning c# as I learn unity, so it's very rough. But I am proud of the hint system I devised.
How can beautiful be semantically close, but gorgeous is cold? Great is close, but excellent is cold?
What the heck do "very" and "great" have to do with the solution, semantically?
https://news.ycombinator.com/threads?id=mwcremer
Also, this post says it’s few hours old, which makes no sense given you commented yesterday. Though the submission history for the user has the correct timestamp too: