Company using Mechanical Turk botches U.S. Senate campaign finance records
publicintegrity.org
publicintegrity.org
Since the topic is public officials acting in their official capacity, I disagree.
I do the technology at an MTurk competitor that automates the quality assurance/training process and pays fair wages to refugees to do the work (workaround.online) if anyone is interested in an alternative to MTurk that is still relatively cheap I would highly recommend it.
I work in a company who have considered it many times for various things, but between the Mechanical Turk API/tech being pretty terrible, and it all being expensive and low quality, we always either end up getting a temp in for a day or two to sit in front of Excel and tidy up data, or if it's a bigger process we outsource it to a data processing company in Bangladesh where we can have dedicated people on our account who sit in a shared Slack channel and who we can train.
What is the current cost per HIT over years previously?
Our outsourcing gives is far better communication and the ability to train staff doing the processing over time, feedback on their performance, and help them get better. I don't know the figures, but I suspect it's a similar price but with far better accuracy, but we do have enough consistent work for this to make sense - if we were more spikey in our demand then it might not.
We regularly ran into these two situations:
- All three workers got different answers
- Two of the three workers agreed on the wrong answer
I think five or more runs may be necessary for data transcription on MTurk.
The error rate I get for data entry tasks is around 0.5%-1% discrepancy between double entry. If you use prior reliability of the worker to tie break between who's right it drops to <0.1% error rate.
But it took forever, and it was relatively expensive for the newspaper. This was way back in the early 1990s. Since then, paper filings tapered off, and electronic filings replaced them. It really is a far better way to do it. Clearly, these records need to be digitized to create public transparency, and that need is apparently being met by bottom-feeding tax-eaters doing a minimum-effort job at a top-shelf price. I am not surprised these records are being botched, but I am surprised this story is coming out only now, rather than back in 2001.
>Reform advocates say this is in large part because of opposition by a small group of Senate Republicans, most notably Senate Majority Leader Mitch McConnell
which is exactly what I'd expect from the dude, but not something I feel like screaming about on HN.
Then Quora came out. Nobody was getting paid but suddenly experts were answering questions in their field.
Wonder if there could be a Quora for MTurk
> Captricity administers this kind of work through Mechanical Turk, an Amazon-owned online labor marketplace.
Ah yes, "groundbreaking collaboration between humans and computers", which in this case actually means "pay a human below minimum wage to type numbers on a keyboard".
Gotta love corporate marketing speak. The only thing that would make this even more ridiculous is if Captricity claimed they were using the mystical "machine learning" too. Hmm, let's check their website. [1]
>Captricity then uses sophisticated machine learning to package up these fields (we call them “shreds”) into quickly identifiable packets.
...lol
1: https://support.captricity.com/blog/captricitys-secret-weapo...
...as in, the results should be inserted into a paper shredder.
You have zero idea what their technology looks like or what it's level of sophistication is, but don't let that stop you from posting this cynical snark.
I don't? I dunno if you missed it, but at the top of this page there is a linked article specifically talking about the sophistication of their technology, and particularly, how it failed to catch even the most egregious errors in their process. Maybe you should read it, it's really interesting.
>The link you provided explains pretty clearly how machines are used, and that is in splicing up the documents and doing first-pass OCR on them.
Yes, and you'll notice that my comment didn't say anything about the usage of the machines, and was focused on the marketing speak phrase of "groundbreaking collaboration between humans and computers" and "machine learning". Regardless of how the machines are being used, those phrases are meaningless bullshit.
>I imagine they're only delegated to humans under a certain confidence level.
Let me get this right: you're criticizing me for not knowing what their technology looks like, and then in the same comment, you yourself are speculating about what their technology looks like. Really?
> Ah yes, "groundbreaking collaboration between humans and computers", which in this case actually means "pay a human below minimum wage to type numbers on a keyboard".
I agree that this is marketing speech (which is a reality of any business selling a product), but you implied that their operation is human-powered with only the illusion of machine interaction, which you go onto confirm:
> Gotta love corporate marketing speak. The only thing that would make this even more ridiculous is if Captricity claimed they were using the mystical "machine learning" too. Hmm, let's check their website.
If you had bothered to read the site you linked (or if you knew what OCR stands for), you'd understand that there is machine learning at play here, and not just for chunking up the documents. You made empty speculations about a company's technology for the sole purpose of denigrating them.
You might think this is splitting hairs, but these sardonic swipes lower the bar of intellectual discourse on this forum. It's OK to speculate about the technology in play, but not when you're insinuating that it's essentially fraudulent with absolutely zero evidence.
Yea, and again, none of that is relevant here. I didn't make any judgements about their actual process (other than the fact that it did, objectively, fail). My focus is on the marketing speak, and despite what you may think, "machine learning" is as useless and as meaningless of a word as "cloud" is.
>but you implied that their operation is human-powered with only the illusion of machine interaction, which you go onto confirm:
I didn't imply any such thing, but that's a nice straw man you're building over there. You know, these arguments like yours are the type of thing that really lower the bar of intellectual discourse on this forum. Wait a minute...
>You might think this is splitting hairs, but these sardonic swipes lower the bar of intellectual discourse on this forum. It's OK to speculate about the technology in play, but not when you're insinuating that it's essentially fraudulent with absolutely zero evidence.
Ohhh okay, so when I make a comment about the technology, it's "lowering the bar of intellectual discourse", but when you speculate about the technology, it's okay? Nice hypocrisy, bud. Talk about lowering the bar of intellectual discourse.
Reminds one of the "Librareome Project" from Rainbows End. That book is dense with the future...
ISTM a ML firm that can't get the ML to work and falls back to Mechanical Turk (how did that wonderful name get past corporate PR?) is kind of like Theranos using normal blood tests to stand in for their imaginary ones: at least as much of a scam on investors as on customers.
Personally, I think the historical reference is clever, but then I'd say the same if someone named an autonomous ship-control system "Flying Dutchman". Maybe I just have bad taste.