Cleaning algorithm finds 20% of errors in major image recognition datasets
deepomatic.com
deepomatic.com
I hope the people in this article had a way to contribute back their improvements, and did so.
It would be so wasteful to have to retrain a dozen models that require a month of GPU time each on to serve as baselines for your new model...
I think it's probably better to have a (say) yearly release of the dataset, with results of some benchmark models released alongside the new version.
This is similar to how Common Voice is handling the problem: it's a crowd sourced, constantly growing dataset, which is awesome if you want to train in as much as possible for production models. You can get the whole current version any time, but they also have releases with a static fileset and train/test split, which should be better for research.
Is it wasteful to throw away a batch of food when 20% of it has been studied to contain the wrong substance, which ends up causing disease?
Isn't it even more wasteful to continue using unedited and unverified data sets just because all the previous models were trained on it, and thus we can no longer advance the state of the research? It's a case of garbage in garbage out.
The comparisons are all relative accuracy, not absolute accuracy. And the comparison is fair. The new technique is receiving the same part-garbage input that the old-techniques were trained on. For the most part, the better technique will still tend to do better unless there's specifically something about it that makes it more sensitive to labeling errors.
And frankly, a percentage of junk has some advantages. Real-word data is a pile of ass, so it's useful for academic models to require robustness.
It seems worrisome that they few percent might be making a coin flip right-randomly instead of wrong-randomly on a mislabelled subgroup of data...
How about XLNet which cost something like $30k-60k to train [1]? GPT-2 may have been around the same [2] is estimated around the same, while thankfully BERT only costs about $7k[3], unless of course you're going to do any new hyperparameter tuning on their models which you of course will do on your own model. Who cares about apples-to-apples comparisons?
We're not talking about spending an extra couple hours and a little money on updated replication. We're talking about an immediate overhead of tens to hundreds of thousands of dollars per new paper.
Tasks are updated over time already to take issues into account, but not continuously as far as I know.
[0] https://www.wired.com/story/deepminds-losses-future-artifici...
[1] https://twitter.com/jekbradbury/status/1143397614093651969
[2] https://news.ycombinator.com/item?id=19402666
[3] https://syncedreview.com/2019/06/27/the-staggering-cost-of-t...
That kind of ruthless experimentation is how AlphaGo was able to exceed even itself. The willingness to say - all these human games we've fed the computer? All these terabytes of data? It's all meaningless! We're going to throw it all away! We will have AlphaGo determine what is good by playing games against itself!
And I bet you that for the next iteration of AlphaGo, the creators of this system will again, delete their own data and retrain when they have a better approach.
If you don't "waste" your existing datasets (once you reallze the flaws in your data sets), you are being held back by the sunk cost principle. You only have yourself to blame when someone does train for the exact same purposes, but with cleaner data.
The person who has the cleanest source of training data will win in deep learning.
You're sabotaging yourself in my opinion. 30k is nothing when you're just sabotaging the training with faulty data.
I don't think you are fully cognizant yet with the formidable scale of AI in the grander scheme of things, as an industry, which is nowadays comparable to transistors circa 1972 in terms of maturity. Long, long ways to go before we sit on "reference" anything. Whether architectures, protocols, models, test standards, it's a Far West as we speak.
You make excellent points in principle, which are important to keep in mind in guiding us all along the way, but now is not the time to set things in stone. More like the opposite.
The matter of the fact is that someone will eventually grab the old and new benchmarks, prove superiority in both, and by that point the new is the one to beat since it would be presumably error-free this time.
- Don't want to deal with vandalism
- Hosting static data is dramatically easier than making a public editing interface
- You want reference versions of the dataset for papers to refer to so that results are comparable. Sometimes this is used as a justification for not fixing completely broken data, like with Fasttext.
https://github.com/facebookresearch/fastText/issues/710
- Building on the previous point, large datasets like this don't play nice with Git. There are lots of "git for data" things but none of them are very mature, and most people don't spend time trying to figure something out.
Imagine if github had an integrated ide for editing large datasets. Also see dolt which is doing good work here.
[1] https://github.com/UniversalDataTool/universal-data-tool
You could maybe split the difference by having an "original" or "reference" version, and a separate moving target that incorporates crowdsourced improvements.
You'd always work with a versioned release when training models, and you'd only typically work with HEAD when you were specifically looking to correct flaws in the data (as the authors in the linked article are).
In general these things are open source, so you can always contribute an improved version of the dataset. But as another commenter said having relatively static ones is also important for benchmarking purposes.
Would also be interesting to see these improved datasets run thru simulation of crashes with existing datasets and see how they handle? Though not sure how you would go about that beyond approaching current providers of such cars for data to work thru and suspect they may be less open to admitting flaws and with that, may be a stumbling block.
Certainly makes you wonder how far we can optimise such datasets to get better results. I know some ML datasets are a case of humans fine tuning and going thru examples and classifying them, and wonder how much that skews or effects error rates as we all know humans error.
It would indeed be very interesting to see the impact of those improved datasets on driving, which is ultimately the task that is automated for cars. We've been working on many projects at Deepomatic not only related to autonomous cars, and we did see some concrete impact of cleaning the datasets beyond performance metrics.
Is that done manually?
Also, do you have strategy for finding errors, where the model learned to mislabel items in order to increase its score? (E.g, red trucks are labeled red cars in both train and test)
This is a great idea if your goal is to maximize the rate at which things you look at turn out to be errors. (On at least one side.)
But it's guaranteed to miss cases where every model makes the same inexplicable-to-the-human-eye mistake, and those cases would appear to be especially interesting.
- you might want to optimize your time and correct as many errors as you can as fast as you can. Using several models will help you ion that case, adn that's actually what we've been focusing on so far.
- you might want to find the most ambiguous cases where you really need to improve your models as those edge cases are the ones causing the problems you have in production. Those 2 objectives are quite opposite. In the first case, you want to find the "easiest" errors, while in the other one, you want to focus on edge cases and you then probably need to look at errors with intermediate scores, where nothing is really sure..
“Annotator agreement” is a measure of confidence in the correctness of labels. And you should always keep an eye out for how these are handled, when reading papers that present a dataset.
Saying we should start doing model agreement is a really good idea imho.
https://en.wikipedia.org/wiki/Active_learning_(machine_learn...
- In the cars-on-the-bridge image, the red bounding box for the semitruck in the oncoming lanes is too small, with its upper bound just above the top of the semi's windshield, ignoring the much taller roof and towed container.
- In the same image, there are red bounding boxes around cars that exist, and also red bounding boxes around non-cars that don't exist. If false positives and false negatives are going to be represented in the same picture, it'd be nice to use different colors for them, so the viewer can tell whether the error was identified correctly or spuriously.
- I have trouble understanding the "bus" screenshot. The caption says "(green pictures are valid errors) – The pink dotted boxes are objects that have not been labelled but that our error spotting algorithm highlighted." In other words, the green-highlighted pictures are false negatives considered from the perspective of the original data set, and the red-highlighted pictures are true negatives. Or alternatively, the green-highlighted pictures are true positives from the perspective of the error-spotting algorithm, and the red-highlighted pictures are false positives. What confuses me is that all 9 pictures are labeled "false positive" by the tabbing at the top of the screenshot.
From memory it had only a small impact (2% strength) with ~7% of results flipped, at 4% it was hard to measure the impact (<1%)
Then the business/political aspects of it, like Tesla demanding somebody who bought a used car pay again for Autopilot.
We already saw crashes by Autopilot users not paying any attention whatsoever (granted AP isn't fully "self-driving", but still).
On top of that, just like with better car safety and even with the introduction safety belt laws, we saw a stark uptick in accidents, that usually affected people outside the car the most, such as pedestrians and bikers. So me being a pedestrian quite often, I dread in particular the semi-self-driving/assisted driving car tech like autopilot, and have a good skepticism when people tell me that the (almost) perfect fully self-driving cars are just around the corner. If my skepticism turns out to be unwarranted, great.
And this tech will keep many consumer cars around longer, in disfavor of public transportation. The one good-ish thing that came out of SARS-CoV-2 is the reduction in air pollution (I am not saying it is a net positive because of that, far from it). The air smells noticeably nicer around here and the noise is also down.
I wish people would stop trotting this one out. Bad actors can deliberately cause humans to crash just as easily if not moreso. If they don't, it's only because such behavior is punishable.
Glitching an AI on the other hand e.g. by holding up a sign is less risky for yourself, and less detectable.
That's not true even allowing for your next constraints, one of which I find to be quite absurd.
In the advanced technological case, you have https://www.theverge.com/2015/7/21/9009213/chrysler-uconnect...
In the non-advanced technological case, you can drop caltrops behind your vehicle as you drive and no one would know it was you.
"But that only happens in cartoons" - Yes, because most people are not cartoon villains. And yet, look, kids throwing rocks, no AI necessary: https://en.wikipedia.org/wiki/2017_Interstate_75_rock-throwi...
> with minor to no risk to ... anybody else
Ah, yes, the ethical murderer who only wants to fuck up just that one car but who sincerely worries about the other drivers on the road. That's the demographic you're concerned about? So how does indiscriminately trying to trick generally available systems specifically target only one person without risking other drivers?
I'll just talk to myself then, because, while I understand you feeling hurt by my comment, I did not attack a strawman.
> Making somebody crash in a dumb car is pretty hard...
Not true. (I gave examples.)
> ...if you want to do it in an undetectable manner...
Still not true. (Same examples.)
> ...with minor to no risk to yourself...
Still not true. (Same examples.)
> ...or anybody else.
Still not true. (This is absurd. Also the same examples still apply.)
Nice ad.
I just see 3 datasets with generic annotations.
"Cleaning algorithm finds 20% of errors in major image recognition datasets" -> "Cleaning algorithm finds errors in 20% of annotations in major image recognitions."
We don't know if the found errors represent 20%, 90% or 2% of the total errors in the dataset.
I'm wondering if those errors are selected on how much they impact the performance?
Anyway, this is probably a much better way of gaining accuracy on the cheap than launching 100+ models for hyperparameter tuning.
It's not necessarily a representation of a better model, but just of a better testing set.
Another example of why you should never mess with the defaults unless strictly necessary.