Even if it happens to be “optimal” in this case at assigning employees to positions based purely on the information available and their likelihood to succeed, biases can present other issues.
It's like having people or neural networks choose door 1 from door 2 without clear advantage to either of them and without making it clear that one door somehow represents "donating blood" while the other represents "kicking puppies".
Yeah, even if the Aima people were 50% better at being a doctor than the Weku (or whatever) we still would not want Aima to be preferred over Weku just for being Aima.
This is the core flaw of this study, imho. The whole equal treatment thing isn't supposed to be "everybody should be equally likely to be picked for a job", but rather "everybody's chances to be picked for a job should only rely on direct characteristics that influence their competence for the job". This study effectively forces the decision maker to use group membership as a proxy for competence due to the lack of information on direct characteristics.
It is hard to see real world situations where there is no performance penalty for structurally choosing participants less fit for the job by using only group membership as a proxy.
Philosophy sometimes says that knowledge is a "justified true belief"*; in this experiment, agents and humans have incorrectly justified a false belief that some applicants are better for certain roles.
* other times, it says this isn't good enough
And then its like you are both saying the justification is incorrect and the belief is false, so its not really like the bare nuance of the concept is adding to the point. Why feel the need to appeal to an (imaginary) authority at all in this case?
"Oh well if philosophy said it, I better be taking this seriously!"
Perhaps a different approach to explain the problem here:
"It ain't what they don't know, it's what they know for sure that just ain't so".
c.f. Sally-Anne test: Sally thinks she knows where her toy is, we know that she doesn't, and indeed couldn't. The LLM (and humans in similar conditions) think they know what the distribution is, we know that they don't.
Only if you're interested in the specific failure modes that LLMs have.
That's all this story is.
Because of that, JTB is often treated as a useful first approximation.
It's easy to see why the JTB criteria are necessary:
Belief: if you don't believe it then you can't count it as knowledge.
Truth: a belief is not (valid) knowledge if it's false.
Justification: accidentally getting the right answer isn't normally knowledge.
But the original claim for JTB was that it was a sufficient definition of knowledge. Later critiques like Gettier's showed that this is not generally true, i.e. there are edge cases for which it fails. In many scenarios, those edge cases don't matter much. So you end up with JTB being an imperfect but useful definition.
Best we can do is have a belief, any belief will unavoidably be justified, but we have no access to any oracle which can tell us if any particular belief was ever "true". When we say we "know" something, we simply have the belief that our certainly is 1, but this is half language and half self-delusion as a certainly of 1 should require infinite Bayesian evidence.