Is That Review a Fake?
nytimes.com
nytimes.com
These are examples of anonymous reviews about hotels that we have in our site:
A positive one => http://www.spottiness.com/spots/BHBZ8QJT
A negative one => http://www.spottiness.com/spots/RKTPXLJJ
Interestingly, they don't have the strong deceptive indicators...
Interesting. Where do "real name" Amazon.com reviews fit in? These make a selling point of being attributable to real people, and to me, imply sincerity, since often these reviewers seem to write reviews almost as a hobby, and often make a point of covering both good and bad aspects of a product.
There's a interesting dynamic here, since Amazon is vanishingly unlikely to harass you on the web, unlike say an ebay seller, who might well come after you if you leave anything other than a perfect review. In this case you are likely to be anonymous as far as everyone but the seller is concerned, yet being sincere may carry some risk to your own ebay account.
Or indicates a lot of fake reviews or something in between.
- Take the top 20 hotels from trip advisor
- Filter out all non-five star reviews, plus any non-english, excessively short, or first-time reviews. Then sample (via a log normal distribution on review length) 20 reviews per hotel. This is their "real" dataset.
- Use Mechanical Turk to collect 400 reviews for these hotels. Turkers are instructed to pretend they work for the marketing dept of the hotel and are to write deceptively fake reviews. Again, quality filters on length, user approval rating, and deduplication are applied. Turkers are paid $1 per review. This creates their "fake" dataset.
I suppose one could still argue that there are selection bias issues here. The sample size is also moderate. Nevertheless, it's a novel approach and you have to start somewhere. Interesting work.
For example, the naive Bayes classifier knows the a priori distribution of review spam (which appears to be held to 50%), but do the undergraduate human judges? It would appear not, given that one judge only labeled 12% deceptive.
Likewise, were the human judges able to see examples of truthful and deceptive reviews before beginning the task? (In other words, are the human judges solving a different problem, e.g., "deception detection", than the classifier e.g., "similarity to prior deceptive reviews from Turkers").
If these are differences between the human and computer annotator setups, are they major differences? Can you spot any other big differences between the two experimental setups?
"Write a good that passes this, this and this filter by a fair margin"
You might try seeing who's hiring Turkers for what. It might give you an idea how much filtering is needed.
The grumpy old man in me wants to suggest that anyone who actually learned how to write would be marked as a fake. Is well-written English so hard to come by nowadays that it makes people suspicious when they see it online?
(Edited to add: apologies if I'm saying something that's covered in the article--couldn't get past the paywall.)
No wonder Google proactively sought out that kid who worked on the paper.
...which is kind of funny if you think about it: to develop a classifier that can identify real reviews, the first thing they do is create a classifier that can identify real reviews to produce a training corpus.
Then they try to approximate the output of their first (rule-based) classifier with a machine learning classifier.
But then you have to grapple with whether or not the bad review is true or not(i.e a competitor posting it). In those cases it's best to focus on places where the reviewer's reputation can be checked.
i.e. all those review sites that let anyone review are more or less worthless. But a blog post by someone with hundreds of posts is a lot more likely to be a real story.
Not too long ago a friend of mine contributed a chapter to a programming book for one of the major publishers and they made sure to tell him to ask his friends to write reviews for the book on amazon, down to detailed instructions how to fake it "properly" and in accordance with how they were trying to position it on the market - like saying that "it is great for beginners to learn X" but one should also add a negative point about how this-and-that chapter needed more clarification or more examples or images.
Also, it's not as if every spammer in the world is aware of this software and will adjust to it accordingly. In fact, I'd say the vast majority wouldn't even be aware.