That is a very poorly formed study precisely because systems like the Mechanical Turk actually work surprisingly well. This is a common theme in the 'wisdom of the masses.' If you ask a single person how many beans are in a jar, you're going to get an answer that's generally very wrong. Yet ask 100 people and average their answer and you tend to get an answer that's oddly
extremely close to the correct answer. As the number of people
independently asked approaches infinity, the error approaches 0.
So when you ask 'x' people to independently judge something, let alone something that is multiple choice with a correct answer, you're going to get answers that are far more accurate than a single individual would give you. So looking at the average answer of 'x' people, comparing it to an AI system, and arguing that they're relatively close so individual people are relatively close to the AI is completely fallacious.
---
As a tangential aside, there's another quirk to the wisdom of the masses. When you let the people communicate and try to intelligently organize and use expertise to come to answer, this effect disappears and the final answer again tends to be very wrong. Kind of an interesting perspective on the current zeitgeist of society and work.
===EDIT===
The authors were obviously aware of the wisdom of the masses. Quoting the paper itself:
To determine whether there is “wisdom in the crowd” (7) (in our case, a small crowd of 20 per subset), participant responses were pooled within each subset using a majority rules criterion. This crowd-based approach yields a prediction accuracy of 67.0%. A one-sided t test reveals that COMPAS is not significantly better than the crowd (P = 0.85).
That's a quite silly p-value and on top of that I'm not sure how they claim their system actually controls for the wisdom of the masses.