Error-riddled data sets are warping our sense of how good AI is
technologyreview.com
technologyreview.com
However, per [1], progress on current ImageNet still correlates with true accuracy. This is in large because we move to harder tests when the easy ones stop working. In part it's also good, because training with label noise forces the use of better-generalizing solutions. The current SOTA is Meta Pseudo Labels[2], which is a particularly clever trick that never even directly exposes the final model to the training data.
[1] Are we done with ImageNet? — https://arxiv.org/abs/2006.07159
[2] Meta Pseudo Labels — https://arxiv.org/abs/2003.10580
Whereas there are major practical issues with even small or subtle errors in, eg., medical or legal datasets for real-world models, a pure research dataset is generally fine as long as a higher score correlates with a smarter model. If MNIST, a nowadays-trivial handwritten digit recognition dataset, had all its ‘1’s labels swapped with ‘3’s labels, it would pretty much fill its role just as well.
This is a good observation because as the saying as "Garbage in garbage out" or "Model is only as good as data". It is better to focus on better data than spending time or improving models or overfitting complicated models to inaccurate data
> The macro-averaged results show that the ratio of correct samples (“C”) ranges from 24% to 87%, with a large variance across the five audited datasets. Particularly severe problems were found in CCAligned and WikiMatrix, with 44 of the 65 languages that we audited for CCAligned containing under 50% correct sentences, and 19 of the 20 in WikiMatrix. In total, 15 of the 205 language specific samples (7.3%) contained not a single correct sentence.
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets — https://arxiv.org/abs/2103.12028
In the face of stuff like that, a 90+% correct dataset doesn't seem like a big deal.
Also, technically accuracy as a metric is robust to noise (https://arxiv.org/abs/2012.04193). That means that a model with the highest accuracy on a noisy dataset will likely be the best model on the clean dataset. So these noisy datasets can still be very useful for the development of deep learning models. In fact if you look at the tradeoffs between getting larger datasets that have noisy labels vs. smaller datasets will clean labels (since good annotation is expensive!), the noisy large-scale dataset will probably be more useful.
Which might be okay depending on the application, but for many data collection processes will not be okay. Imagine if one group of people's wealth is systematically under-estimated (aka "measured") when training an insurance policy AI. The algorithm will then correctly learn the bias in the dataset.
None of this is magic. Data collection is hard. Spotting your own biases before they make it into your dataset is a blind-spot exercise.
You mention an "insurance policy AI". The rest of this thread seems to be about image recognition. What's an insurance policy AI and how does it work? Is it a thing or a hypothetical future thing?
Can bias effects have equally bad effects in image recognition? I know about the story of the photo software that categorized a man's holiday photos under "gorillas" because it was trained only on white people. This is terribly insulting, but less terrible than an insurance company unfairly overcharging you or refusing you service.
I guess what I'm saying is I can't come up with biased image recognition AIs having unfair outcomes that actually affect people deeply, and I'm likely missing lots of terrible examples, and I'd love if someone can educate me.
EDIT: before people erupt in outrage, I'm not saying that a computer telling you that you're a gorilla isn't terrible, but in the end it's an indictment of the software, not you.
It may be one of the precursors to intelligence: The ability to cognitively filter out unnecessary sensory stimulus in order to focus perception to the object or subject of a creature's attention.
That's exactly the opposite of what this article says.
The observation that a less faulty model is likely less accurate on a noisy validation set than a more faulty model, doesn't change the fact that there must be faulty models with higher accuracy than a perfect model on a noisy validation set.
Edit: I said "surely", but, no: GoogleNet's paper has no such section: https://static.googleusercontent.com/media/research.google.c...
Researchers have no incentive to refute their own success; and neither do their peers who are producing similar "research" themselves.
The whole game is to state enough metaphorical propositions until they seem concrete.
"New AI interprets conversations and writes stories!" they say...
Their evidence? Run it enough times until a subsample appears to match our expectations, and show that subsample.
What about all the other counter-examples? That's compared with "errors humans make too!!!!"
When, of course, if you look at those counter-examples they completely destroy the interpretation of the system as understanding anything.
It is ruled out by the data analysis procedure: measurement data is fed to algorithms which assume it is sampled randomly. The whole purpose of "intelligence" (, understanding, etc.) is to uncover the non-random underlying model which explains the data distribution. Thus statistical AI simply cannot do what it is claimed of it. All we are left to do, for marketing and financial reasons, is play a game of metaphors and superstition.
Personally I find it infuriating. It's profoundly religious/superstitious and a debasement to scientific practice: of course, computer scientists aren't scientists; and here is where the problem arises. They have no empirical sense of what a model of the world actually is; nor of what a "mere statistics of measurement can do".
A supervisor of mine tried once to make me switch to more AI-based thing by telling me my research field was niche, and asking if "AI in education" wasn’t something bigger and more interesting (I replied no). In retrospect I should have asked him the question back by rephrasing it like proposed above.
Especially these days, when human computers have basically disappeared as a job.
https://labelerrors.com/static/imagenet/val/n03000134/ILSVRC...
It is also interesting that their model guessed "container ship" with the dominant element of the picture being a harbor crane. I think there might be a lot of images in the ImageNet data set that include both a crane and a ship as strong elements, but only have the ship label.
ImageNet is a single label set, not a multi label set and that is a known problem. This paper talks more about it: https://github.com/naver-ai/relabel_imagenet
With large data sets containing millions of images the only way forward seems to have AI keep refining the data, escalating things multiple automatons disagree on to be reviewed by more classifiers, including humans.
* The label is for something in the image, but there are multiple things in the image.
* Getting animal species wrong.
For some reason though I've never heard of "robust ML" and I always wonder why. The only thing I hear about is increasing NN sizes and increasing model complexity. Many smart people are working on this so I suppose there's a reason for that but it's just not very obvious to me. Can someone with knowledge of the matter provide some insight?
Neural nets typically don’t benefit much from it because you can use batch normalization, dropout and clever activation functions to achieve the same results, by having the network learn diminished sensitivity to outliers that produce neurons which saturate the low end of an activation function.
This is preferable because many of the robust potential functions involve absolute values, order statistics and other non-differentiable quantities that are hard to put into backpropagation-based optimizers. You almost always would need to relax the loss function to something that trades off smoothness against outlier robustness, where convergence will be slower and slower as you crank the trade off closer to outlier robustness.
So long as AI best practice is pretending to pay, will pretending to work will be the outcome.
At best.
Maybe I would find this more interesting if the question were, "how good does good enough have to be"?
Where you really want/need accurate labels is the eval set, which the article addresses.
Data QA is just not sexy, and doesn't get you promotions, or the big paper.
Most politically interested people who aren't techies never really understood a lot of really important ideas from the "technology, power and freedom" section of the library. Consider these titles:
- "Why Linux, Wikipedia & WWW are existence proof
for anarcho-communism/-capitalism"
- "Youtube & Twitter are protocol squatters"
- "RSS podcasting, the last free online media"
These don't mean anything to most people, and it's hard to get them to care. "Net Neutrality" did make its way to politics, but in a much bastardized form and most people never understood it.. even reporters and such.OTOH, the political and economic of datasets is instinctively understood by non-tech people. The average person (certainly reporter) is highly in-tune to the political importance of imagenet's Categories for Bad Person. Ownership of the datasets. Control over contents. etc. Often ahead of tech-oriented people.
The person who struggled understanding "free as in speech" applied to software sounds like Stallman on chivas when the conversation is about datasets. Interesting.