Label a Dataset with a Few Lines of Code
eric-landau.medium.com
eric-landau.medium.com
It works well if a domain expert can say something without “cheating” and looking at the data like “put a box around round red objects because those are always apples”. But in practice people tend to cheat and look at the data first, and you end up with humans trying to emulate ML, poorly.
A more varied dataset will require additional strategies. We have done this type of thing with various datasets and what normally works is a combination of some vertical models, heuristics specific to the dataset, classical computer vision techniques, and some human label seeding/correction.
-The outputs of the auto-labeler. If this is strong, you've learned that you didn't need the training set after all - you managed to solve the problem without it!
-The outputs of a model trained on auto-labeled data. If this is strong but the above test was not, then this pipeline makes sense.
-The outputs of a model trained on human-labeled data. If this is strong but the above tests were not, we're in trouble.
If none of the three are strong, then the training data was lacking (assuming we've done our best on tuning the model we're trying to train), and so no real value was gained by annotating it.
When you are labelling data, you have access to strategies and means that might not be available to your downstream model. In our experience this includes a human in the loop component, building non-robust ensemble models(we call these micro-models), and some "guess work" functions on the data. All of this together can make an "auto labeller" that does pretty well getting labels made, but really the sum of these strategies is very different from some well trained neural network that will be running on edge or whatever.
The point of a model is not to label the data, it's to generate some value in some out of sample task, quite different from strategies that you can run in a sandboxed environment with your training data.
Obviously we should expect that the auto-labeler fails on the test set, because we assume we're exploiting some convenience that won't be available at test time. But we should still try - it might reveal that our task is too easy to need the model we were planning to train, or it might reveal that our test set is not actually representative.
So it comes down to how good the auto generated labels are(from a human perspective), which is a fair point that I didn't address much in the article, but in general comes down to a good QA process(which is applied to both human labels and machine labels equally because humans also make mistakes in this stuff).
In the article the dataset was small enough and the labels simple enough that I could run very quick visual inspection over the results, but for more complicated tasks we have a more rigorous human review process for evaluating label accuracy(again to both human and algorithm produced labels). The auto generated labels may not be more efficient overall if they require a lot of correction after review, but for this case, and a lot of other ones, they just are empirically are.
However the thing is about labels being noisy and using multiple labeling strategies to help train a higher fidelity model - https://www.snorkel.org/
You are right there are a bunch of difficult problems this technique isn't perfect for, but it actually can still help improve the efficiency of labelling a lot and when I do it I get the added bonus of understanding the dataset a lot better.
Otherwise I overall agree with you that we should consider this where we can .. as evidenced by my snorkel link.
I think we are lazily giving up our intellectual power to models hoping that they will just discover patterns by magic, where it is actually very worth to go through the data science process starting with labelling because we actually learn as humans. The thesis is that this will also make our DL models better in the long run. We would never have come up with cool algorithms if we just always outsourced this work to models.