Wait... the naive-Bayes was trained on Yelp data? Isn't Yelp data also crowd-sourced information? I may not be thinking about this right, but it seems to me that training the classifier on crowd-sourced data and then comparing that to Mechanical Turk... that in the end you're just comparing the quality of the crowd-sourced data to each other?