- Take the top 20 hotels from trip advisor
- Filter out all non-five star reviews, plus any non-english, excessively short, or first-time reviews. Then sample (via a log normal distribution on review length) 20 reviews per hotel. This is their "real" dataset.
- Use Mechanical Turk to collect 400 reviews for these hotels. Turkers are instructed to pretend they work for the marketing dept of the hotel and are to write deceptively fake reviews. Again, quality filters on length, user approval rating, and deduplication are applied. Turkers are paid $1 per review. This creates their "fake" dataset.
I suppose one could still argue that there are selection bias issues here. The sample size is also moderate. Nevertheless, it's a novel approach and you have to start somewhere. Interesting work.
For example, the naive Bayes classifier knows the a priori distribution of review spam (which appears to be held to 50%), but do the undergraduate human judges? It would appear not, given that one judge only labeled 12% deceptive.
Likewise, were the human judges able to see examples of truthful and deceptive reviews before beginning the task? (In other words, are the human judges solving a different problem, e.g., "deception detection", than the classifier e.g., "similarity to prior deceptive reviews from Turkers").
If these are differences between the human and computer annotator setups, are they major differences? Can you spot any other big differences between the two experimental setups?
"Write a good that passes this, this and this filter by a fair margin"
You might try seeing who's hiring Turkers for what. It might give you an idea how much filtering is needed.