We're talking about Amazon, one of the biggest powerhouse ML employers. I don't buy that the model was poorly designed or ineffective. They also didn't just scrap the model without understanding how or why it failed to meet its objectives.
And whatever the cause was, it was not the poor quality of the training data. They tried to stop the model from downranking women based on obvious keywords, only to find it learning to downrank them based on more subtle language cues:
> Amazon edited the programs to make them neutral to these particular terms. But that was no guarantee that the machines would not devise other ways of sorting candidates that could prove discriminatory, the people said.
So the answer is 3 or 4.
If the answer was 4 then they would have probably mentioned the cause of the bias somewhere in that otherwise detailed article. But they didn't, possibly because the cause is controversial - probably option 3 but possibly still option 4.
And then there's the subtle cop-out:
> Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs, the people said. With the technology returning results almost at random, Amazon shut down the project, they said.
If the model was actually useless and returning random noise, then there wouldn't be any bias, and the article wouldn't need to talk about discrimination. This paragraph reads to me like they decided to mention long-tail results (that you'd find in any ML model) as supportive 'evidence' that the model was somehow broken rather than producing valid but controversial results.