Show HN: Sense2vec model trained on all 2015 Reddit comments
sense2vec.spacy.io
sense2vec.spacy.io
However, I'm not sure if I agree with the use of phrase similarity as a indicator of Reddit hivemind behavior, which was discussed in the complementary blog post. The tool is more of an indicator of the writing styles of Reddit's primary demographics (Male, 18-30) and phrases which are coincident, instead of weighting the importance of given phrases to Reddit discussion.
If you also trained a model on a different news aggregator's comments, would it be possible to "match up" the meaning vector spaces and see differences in what meanings each community ascribes to words?
Additionally, could you determine sentiments from the positions of the words in the meaning space? One example, looking for underlying assumptions, like associations of the words 'black person' and 'criminal'. Or idk, man + sex = player, but woman + sex = slut.
Would it be possible to go higher level and see how much a corpus agrees with "free markets are good" based on its word positions?
It seems like word2vec has the potential to bring Sapir-Whorf to a whole new level.
What you need is a 'seam' that connects the two vector spaces. I would do it like this.
Train a single model, with the words decorated by their subreddit. So you have combat:/r/gaming and combat:/r/history as different tokens. Then you have shared tokens which aren't decorated in this way.
http://bookworm.benschmidt.org/posts/2015-10-30-rejecting-th...
https://github.com/spacy-io/sense2vec
For installation:
$ pip install -e git+git://github.com/spacy-io/sense2vec.git#egg=sense2vec
$ python -m sense2vec.download
This downloads and installs the sense2vec package and model we used in the demo.
We are currently working on finishing the PyPI package. Meanwhile only Linux is supported and the docs are pretty much non-existing. Also, please make sure you have a recent Blas/Atlas package installed (RedHat: atlas, atlas-devel)
Direct link to the model file (~600MB):
https://index.spacy.io/models/reddit_vectors-1.0.1/archive.g...
something like this definitely makes word net obsolete. :-( http://wordnetweb.princeton.edu/perl/webwn
But there is still room for improvement. Searching "Haskell" leads to Clojure and C++. Yes, these are both Programming Languages, but out of all I personally wouldn't have said C++. :)
"Scheme" leads to Haskell, witch is very fitting, and Prolog, witch seems to fit as a "university language" aswell.
For "Agda" it outputs "typeclasses" at 74%. This is a much discussed topic. But for a "truly" semantic understand It should know that both "Agda" and "Haskell" are in the same category and that "typeclasses" is a property that elments in this category have or don't have.
Still, very impressive. But not the singularity jet.
(first name, not first word, had to scroll down a bit, past words like 'Nazi soldier' and 'evildoer')
The last line (Parent commenter can toggle NSFW or delete...) has this formatting:
^Parent ^commenter ^can [^toggle ^NSFW](/message/compose?to=autowikibot&subject=AutoWikibot NSFW toggle&message=%2Btoggle-nsfw+ct3omf8) ^or[](#or) [^delete](/message/compose?to=autowikibot&subject=AutoWikibot Deletion&message=%2Bdelete+ct3omf8)^. ^Will ^also ^delete ^on ^comment ^score ^of ^-1 ^or ^less. ^| [^(FAQs)](/r/autowikibot/wiki/index) ^| [^Mods](/r/autowikibot/comments/1x013o/for_moderators_switches_commands_and_css/) ^| [^Call ^Me](/r/autowikibot/comments/1ux484/ask_wikibot/)
You can see how that formatting might fuck up their nice natural language processor and tokenizer.I've been working for a while on extracting "semantic" from naked text (mostly news). One of the big limits of word2vec is that the semantic should be related with words proximity in sentences, which works in some cases, but not always.
I'll give an example for all: travel/geography. If you query something like Italy [1], the results are other European states. But if you're looking for news in Italy, or to plan a vacation to Italy, or for some Italian food... or anything related to Italy itself, you probably don't expect "Spain" to be the first result.
It would be nice to have some sort of easy way in word2vec to define domains and their relationship with words proximity in sentences, to overcome situations like this one.
One way would be to predefine the entity types or tags that you want to get in your results. So you could ask for things like Italy that are nouns.
The other way is to use the vector space. The classic demonstration of this is the arithmetic, doing like "Italy|GPE - *|GPE + food". My results for this have been very mixed. I wouldn't expect the query above to work.
I would think you'd have more luck specifying the query as a combination of constraints: first query for foods in some way, and then sort them by distance from Italy.
It's more like a recommender: If I like Pizza in Napoli where should I go in Spain and what should I eat there?
Italy:Pizza -> Spain:?
Italy:Napoli -> Spain:?
Used POS tagging in a previous post, though not with Word2Vec since I wasn't sure if differences like duck verb vs. duck noun would improve the result because the placement of verb vs. noun would already be different. Though certainly an interesting approach, I'm wondering if going backwards might yield better results for the POS tagger as well, since verb vs noun would span disparate word clusters.
1) http://dbunker.github.io/2016/01/05/spark-word2vec-on-reddit...
Sometimes I get tricked when I enter a query that has the same top result as the one that's currently displayed. Then it looks like the results haven't changed, but further down the list, they have.
Edit: We have this fixed but we're reluctant to roll it out. Two days ago when we did an AWS deploy, they replaced healthy machines with unhealthy ones, and we had an hour of outage.
If you want to fix this on the client, I think the following quick fix works. At the bottom of the sense2vec script https://sense2vec.spacy.io/js/sense2vec.js , change:
input.addEventListener('keydown', function() {
if(event.keyCode == 13) run();
});
to input.addEventListener('keydown', function(event) {
if(event.keyCode == 13) run();
});It also does not find anything for "duck sized horses" or "horse sized duck".
It seems to have learned the relationship between characters and shows, though it always ranks the Monogatari characters as relatively close to any character search.
I searched "dank memes" and got "steel beems". Other searches all failed.
edit: and now it doesn't?
First result is GOP
Clearly "Obama" isn't an organization, but it does a decent job. Consider the sentence "Obama’s unsolicited advice that Congress should “go ahead and vote” has only hardened the GOP majority’s resistance."