Stanford researchers to open source model they say has nailed sentiment analysis
gigaom.com
gigaom.com
Ridiculous accuracy for something as complex as sentiment analysis. You don't hear established researchers say something like this often. Moving any number from 85% to 95% is the work of the gods.
I wonder if the code they release will include some version of the data from the mechnical turk project. Code for this is great and many people (myself included) will be able to learn a lot from it. But it won't have the same level of reproducibility without the data.
If they do release it they will effectively be giving away the money they spent on mechanical turk. 11,000 HITS ain't cheap and they probably had redundant sampling. If they decide to make this data public as well it would be a big win for research because labelled data is so important to machine learning work.
Open sourcing the code associated with a research paper is already a huge deal. It's great to see big name researchers like Andrew Ng pushing the trend for publishing code. If nothing else this is a great example for computer science papers going forward.
I don't know if the "raw" data from MTurk is included in the data set. But at least the finalised data has been released for quite some time now.
http://nlp.stanford.edu/~socherr/stanfordSentimentTreebank.z...
As for the value of the data, I have heard numbers around 10,000 USD. But I have little more than academic gossip to back these numbers up.
I haven't read the research but I'm assuming he used a fairly unrestricted crowd. I'd be interested to use the crowd to rate sentiment analysis results from this model vs. CrowdFlower's Senti to see how the model fares against a specialized sentiment analysis crowd. I would be very surprised if the model won. I am willing to run a comparison if there is sufficient interest in the data.
Sentiment analysis is often akin to mind reading. I don't know how the average human would do, but given how often people miss jokes & sarcasm on forums & emails it wouldn't surprise me at all to learn that the average human analysis would score below 85% on single sentences without context.
(And to be fair to the OP, they said 95% would be the work of gods, and I'd certainly agree that would be better than most humans could do)
http://engineering.stanford.edu/news/stanford-algorithm-anal...
This is really fun to play with, and I'm surprised how well it can parse the sentiment of sample sentences I threw at it. I've tried a couple random examples (like "I don't know what the artist was smoking, but the song made no sense (though I liked the beat!)") and have not yet gotten a wrong analysis. Even the phrase parsing is pretty spot-on.
As a side note, this is much more interesting than the "sediment" analysis I excepted after skimming the title. (Unfortunately though, the analyzer got this final sentence wrong: http://cl.ly/image/301u1q46263m)
Edit: seems like this system could get significantly more robust with more data. If you look in the comments section, you can see some comments from the professor himself, i.e. "Possibly because the word "buying", only appears once in the entire dataset and it's in a pretty negative context: http://nlp.stanford.edu/sentiment/treebank.html?w=buying"
If you gave it 100,000 phrases, I wouldn't be surprised if it could hit the 95% mark that Socher mentions.
I tried to fool it, but it took quite a bit of effort to contrive a sentence that it got wrong.
The sentence I finally fooled it with:
--> I enjoy clipping my toenails to paying any more attention to this movie.
which it rated as highly positive.
Still, I'm not even sure if using the preposition "to" in that manner is proper english.
>I'd enjoy clipping my toenails more than paying any more attention to this movie. --> positive
>I'd enjoy clipping my toenails over paying any more attention to this movie. --> negative
If you change "paying" to "giving" it classes that clause as neutral instead of negative, but doesn't change the result.
is parsed as being positive ... I guess the sarcasm detector still needs some work :)
>It is, without question, the worst film ever made. (--) But this comment is in no way meant to be discouraging. (-) Because while The Room is the worst movie ever made it is also the greatest way to spend a blisteringly fast 100 minutes in the dark. (--) Simply put, 'The Room' will change your life. (0)
[1] http://nlp.stanford.edu/~socherr/EMNLP2013_RNTN.pdf
Edit: formatting
One issue is that a system like this needs to be trained for the specific kinds of documents you are processing. For instance, if you are looking about people's opinions on stocks, there is specific terminology to look for such as "buy", "sell", "short" or "long", "missed earnings", price targets, etc.
This isn't so much a problem with their method, but it is a problem w/ the specific model they are publishing.
I like that they are using "beyond bag of words" methods and I find it very believable they could get much better results if they had a bigger training set and more effort in tuning.
One advantage us commercial folks have is that we don't need to bet on every hand. Reviews like that one of the "Room" are ambiguous at best and should be filed as so.
"Sentiment analysis" is too broad of a category to really cover in a single article like this. What they've done is taken a very difficult problem, sentence-level binary sentiment, and made solid progress on it. The baseline for this dataset using totally naive techniques is around 75%, and their results are the state of the art.
The move from 85% to 95% isn't really an interesting one. What really matters is exploring the numerous other open questions in the field of affect recognition, notably two thing:
* Sentiment at different granularities. Document level analysis has been far above 90% for years; this work is pushing forward sentence level. Other work is making great progress on targeted opinions even finer-grained than that, like looking at specific attributes of products. What if you like a movie's acting but not its plot? This structured nuance is not addressed here.
* Domain adaptation. You talk about movies in a different way from almost anything else. A movie review is positive if it's unpredictable; your opinion of the unpredictability of dishwashers or political candidates is probably different. For anything beyond movie reviews this method may work, but this particular dataset certainly won't.
Looking forward to seeing more from this group, as ever; Chris Manning's research team has an excellent reputation in the field.
Can someone here please explain whether the use of Mechanical Turk here is a cop-out from building a better computational model, or just an ordinary use of supervised learning in place of unsupervised?
- certain combinations of words within phrases score all over the place
- hand those to mechanical turk for human classification
- understand where the results differ from the model
- patch the model where necessary when it breaks down.
The example they gave with the "but..." at the apex of the sentence is difficult primarily because it's ambiguous to what proceeds it. It could be positive or could be negative, especially from a programmatic standpoint.
Really fascinating stuff. Can't wait to see the code.
If the model is telling us that it is uncertain about specific examples, and would like more information on examples like those, that's active learning.
That sounds different from what you describe in your post, depending on what you mean by 'sometimes the model is wrong, so we hand the data off to a human being'.
It depends on how we know the model is wrong.
If we know its wrong on a test datum, which is part of a big set of test data humans labelled without any input from the model, then its standard 'supervised learning'.
If, instead, the model is 'wrong' because it expresses uncertainty for particular test data, then, if we go and have a human classify that data it was uncertain about, and retrain the model, then we are probably doing Active Learning. In this case, the model/system is (at least partly) guiding the learning process.
Reinforcement learning is neither of these things exactly - it describes a more general framework, where the system is getting rewarded based on how well its performing.
Lets say you want to choose 1 of 5 labels for each datum. In supervised learning, the system gets given the right label for each training example. In a RL setup, it might be shown an example, have to guess a label, and maybe be told if it got the right guess, but if it guessed wrong, just told it was wrong - but not necessarily told what the right answer was.
There's a little fuzziness to how all these terms are used in practice.
[0] http://en.wikipedia.org/wiki/Active_learning_(machine_learni...
So far I think they're accurate. Homework 1's due in a few days, but the lowest 2 homeworks are dropped.
The talk they gave, "Deep learning for NLP (without magic)" was pretty good: http://techtalks.tv/events/312/573/
You can find more details here -http://arxiv.org/ftp/arxiv/papers/1305/1305.6143.pdf and the code over here - https://github.com/vivekn/sentiment/blob/master/info.py
> The movie has been signed by Michael Bay: this is the same man who directed "The Rock" in 1996; now he has made "Transformers: Revenge of the Fallen", and, well, Faust made a better deal.
It correctly identifies the sentence as negative, while all words taken individually are either neutral or positive... I'm impressed.
As I'm reading though the article I see that it says the algorithm can understand "Human Language." By this I'm guessing they mean English. One thing I learned about sentiment analysis is that analyzing other languages may prove to be a bit more difficult.
Another question I have is to run it up against this very basic sentiment analysis engine that my old manager built which basically had 13 positive words and 13 negative words and was about 80% accurate as well: no neural networks, AI or machine learning needed.
It would be interesting to read the paper to find out what accuracy really means here. I doubt that human readers agree on the sentiment of movie reviews 95% of the time.
"We’re actually able to put whole sentences and longer phrases into vector spaces without ignoring the order of the words."
Wait, didn't Mikolov et al. (Google) [just figure out][1] how to put entire languages into vector spaces?
"That makes about as much sense as a whale and a dolphin getting it on."
Keep working on it guys... I wish I understood sentiment trees well enough to be able to train it properly for this statement... Is a sentiment tree able to properly represent sarcasm and innuendo? <--- Honest question
I'd say that if you posed that statement to 1000 English speakers from around the world, at least 1% of them would be baffled by it.
All that is to say non-conventional uses of language are always going to be a problem for natural language processing. If a certain kind of innuendo and sarcasm is represented often in its training data, then the model SHOULD be able to understand it when it sees it again.
I also wonder about sentences which could be understood and defended as being positive to one human reader and negative to another.
"That is the craziest thing I've ever heard." or simply "That is sick."
Off the top of my head, I know of a company that's trying to tackle online complaints (VozDirecta.com), another that feeds "what they're saying about your company"...
"I loved this movie.. NOT!"
and it classified it as positive. :)
It also doesn't seem to like swear words and web abbreviations (like j/k for example). Perhaps with more data (from twitter or similar) it could go a long way though; as it definitely needs to learn about the more uncouth words.