CrowdFlower maintained statistics about the accuracy of individual workers. Additionally, they made it easy to include gold-standard "test" questions, which weeds out workers who are not doing the specific task correctly.
Without this sort of quality-control platform, mechanical turk is just unreliable crowdsourced-labor infrastructure. I understand that AWS doesn't want to build too much on the services they provide but, really, mechanical turk sucks.
I'm working with a researcher who is conducting experiments using turk, and he has a half-written shoddy version of what crowdflower offered. It seems that everyone using turk has to reinvent the wheel. Why can't someone provide decent quality control over turk results?