AI Defeats the Hivemind
technologyreview.com
technologyreview.com
We actually used turking at my company for some really nutty stuff, logo generation. Basically we'd give people a URL and ask them to generate a 160x40 logo for it. We had some base rules, like the background had to be solid, have no scaling artifacts etc..
We assigned each logo to five people.
Our reward was essentially this: - anybody who met all the rules, got .25c - the best of all that met the rules, got a 50c bonus
It took a few days for people to get the hang of it, but after that we consistently got excellent results, with some really creative stuff coming back. Yes, we were paying up to $1.50 for the logos, but we weren't using them for every site, only the really popular ones, and having it automated made it worth it. Every day we spent maybe 60 seconds picking the best logo of five submissions for a few dozen sites, everything else was automated.
The product that used these by the way is NewsRoom, a pretty sexy RSS reader available on Android. All the logos you see for sites there were generated by Turkers.
Anyways, finding the right equation for that task took some experimentation, but I was impressed by the results in the end.
79 passed. This was an extremely basic multiple choice test.
It makes one wonder how the other 4,581 were smart enough to
operate a web browser in the first place.
I stopped reading right there.As for the question itself, that's simple: people come for the money, and since "Turkers" are paid pennies for those tasks that means they have to do a lot of them; so replying randomly on a test is a no-brainer (I wouldn't even bother to click and type and just write a script).
It's a good thing we've got these magazines reminding us how we are so smart and the rest of the world is so stupid. What would I do without my over-inflated ego?
At the price per HIT I paid (0.08 to 0.15), I only had to throw out 2-3 answers out of ~2000 due to someone trying to reply randomly.
My own father is "not smart enough" to operate a browser. Lack of English skills don't help him. But he can read French and Russian just fine, he has a Ph.D in his profession and a carrier in politics (former advisor to the prime minister, currently a senator in a eastern-European country).
Also, the paper says 1658 passed, but probably only 79 passed with > 90% accuracy?
I looked at the cited paper and did not see the cost, but without the cost I really would not bother interpreting these results. "Machines work for electricity; humans need real money. News at 11."
EDIT: maybe we can get the real MTers to do the algorithm/problem matching bit...
But then I realized that, in most offices, work done is meaningless next to number of hours spent in the building ;)
Still, if you were motivated, you could get paid to sit and browse BoingBoing all day. :-)
There are a lot of much easier ways to achieve this. As if you'd want to...
From the summary:
The results weren't pretty: in order to find a population of Turkers whose work was passable, the researchers first used Mechanical Turk to administer a test to 4,660 applicants. It was a multiple choice test to determine whether or not a Turker could identify the correct category for a business (Restaurant, Shopping, etc.) and verify, via its official website or by phone, its correct phone number and address.
79 passed. This was an extremely basic multiple choice test. It makes one wonder how the other 4,581 were smart enough to operate a web browser in the first place.
From the paper:
Of the 4,660 workers who took this test, only 1,658 (35.6%) workers earned a passing score, and over 25% of workers answered fewer than half of the questions correctly.
To investigate the high failure rate, we conversed with workers directly on TurkerNation and through private email. Based upon worker’s names and email addresses, we believe that we conversed with a representative sample of workers both inside and outside the United States. We found that the test was not too difficult and that most workers comprehended the questions. We believe that many applicants simply try to gain access to tasks as quickly as possible and do not actually put care into completing the test.
ie, 1658/4660 workers passed this test, NOT 79 (!!)
Then later they describe some additional filtering they put in place to attempt to find the best workers (they tried estimated location and time to complete task). Based on these filters they said: Using a combination of pre-screening and the test tasks described above, only 79 workers of 4,660 applicants qualified to process real business changes.
They are also fairly robust and work well for a wide variety of problem sets.
Other techniques sometimes offer some improvement, but often don't. Generally a Bayesian classifier is a good place to start.
HN submission of the same (13 days ago) here: http://news.ycombinator.com/item?id=1984130