Haha, well, you know how it is. Everyone has their own opinion as to how to solve problems like this. Some people see their role as sort of fire alarms to say "Hey! This thing is happening!". You know, like how every thing has people at all stages in the process. Others might believe that this is one instance (the benchmarks being 'wrong') and that a better outcome would be that the people putting out this stuff are aware of ways that things can go wrong so they think about it in future (maybe this time it's these benchmarks, next time it's the other ones, and then maybe next time it's something inherent to the structure, whatever, it's a bit whack-a-mole to solve at the instance level).
Like, for instance, HN has people who will bring up privacy violations of big tech constantly. They see their role as making sure the conversation is happening. Not justifying this. Just aiming to understand it.
For my part, I prefer to take the approach you're talking about because I, too, think that the fastest path to this is getting the photos, labeling the photos, and then lobbying for inclusion. Ultimately, I think it's okay if things optimize fast for growth and then we fix up issues afterwards. So the people building the benchmark sets weren't able to get a set that's representative of humanity. Should they have waited till they could have done that? IMHO, no. Rapid release moves the state of the art forward and then we can put in all of these corrections as we move.
Then there's the question of whether all-humans dataset is a good thing or if instead a thing that is white-humans and another that is black-humans is better. Anyway, all said, my personal approach to this problem (if I cared about it a lot, which I don't) would be to say "Current benchmarks and training data available bias towards certain races. I'd like to build X/Y to solve that. Here's what I have so far" etc. etc. I think positive engagement like that yields better results because the vast majority of scientists actually aren't weird race supremacists and the vast majority of AI researchers will gobble up any more data you give them which is segmented and labeled differently, the greedy bastards :D
Some challenges that I can still see:
* Getting the data. Might not actually exist.
* Labeling the data. Probably needs some work.
* Getting it into the benchmarks. It'll invalidate old scores, so there's just a product adoption problem here. I don't know how the community handles newer releases.
As a last aside, I suspect that this conversation ended up the way it did because:
a. It's charged. It's race-based differing outcomes. That's a sensitive subject.
b. People feel unheard. This is natural. Like, this is not an 'interesting' problem. It's literally just a data error so the luminaries in the techniques part of the field aren't really that interested in it. And the techniques part is where the sexy is.
c. This sort of thing has a tendency to escalate. One side says "You're not listening to what I say" and the other side says "I'm not racist. I don't get why you're calling me that" and before you know it it becomes "You have to be racist to be ignoring me" and whatnot and de-escalation becomes impossible. Especially because everyone rewards the loudest on each side.
Honestly, I think it's quite interesting to observe and to understand as just a view into the human condition but we use AI models professionally in the GIS space and professionally we just don't go near this at all. No part of me finds it interesting to solve or to interact with the discussion in any way. I only sort of participated in this here because I think I managed some insight into what it is and I wanted to write that down because I wish someone else could have accelerated me into it.
Anyway, I think that's all the insight I have on the subject, so I'm going to just leave it there. Any more and I'll be ass-pulling.