- Take the top 20 hotels from trip advisor
- Filter out all non-five star reviews, plus any non-english, excessively short, or first-time reviews. Then sample (via a log normal distribution on review length) 20 reviews per hotel. This is their "real" dataset.
- Use Mechanical Turk to collect 400 reviews for these hotels. Turkers are instructed to pretend they work for the marketing dept of the hotel and are to write deceptively fake reviews. Again, quality filters on length, user approval rating, and deduplication are applied. Turkers are paid $1 per review. This creates their "fake" dataset.
I suppose one could still argue that there are selection bias issues here. The sample size is also moderate. Nevertheless, it's a novel approach and you have to start somewhere. Interesting work.