A better way to build ML: Why you should be using Active Learning
humanloop.com
humanloop.com
The pros working in big shops who write these tend to overlook the tiny use cases such as apps that recognize a cat coming through a cat door (as opposed to a raccoon) which can get by with minuscule training.
There's a lot of discussion of "big data" but small data is amazingly powerful too. I wish there was more bridging of these two worlds — to have tools that deal with the needs of small data, without the assumption that training a model takes days or months, and on the other side, to have the big data world share more insights about how they manage their data for the big cases. There is a ton of info out there but what I find lacking is info about how labeling and tagging is managed on a large scale (I'm interested in both, big and small, as well as medium). Maybe I'm just missing something. This article gave some good clues — thanks!
Early days but we’re looking to onboard a few more customers to help guide our roadmap.
You can oftentimes do surprisingly well on a smaller dataset with proper augmentation. Unfortunately augmentation techniques don't seem to get quite as much attention as they deserve, perhaps because they're perceived in ML conferences as problem-specific hacks rather than general-purpose techniques.
In one of our research projects, we used AL to improve part-of-speech prediction, inspired by work by Rehbein and Ruppenhofer, e.g. https://www.aclweb.org/anthology/P17-1107/
Our data base was a corpus of Scientific English from 17th-now and for our data and situation, we found that choosing the right tool/model and having the right training data were the most important things. Once that was in place, active learning did not, unfortunately, add that much. For different tools/settings, we got about +/-0.2% in accuracy for checking 200k tokens and only correcting 400 of them.
Maybe one problem was that AL was only triggered when a majority vote was inconclusive. Also, we used it on top of individualised, gs training data. I guess things can look different if you don't have a gs to start with. And if you have better computational resources: Our oracles spent quite some time waiting, which is why we even reorganised the original design to then process batches of corrections.
As so often, those null results were hard to publish :|
Either way, I thought I'd share our experiences. Your work sounds really cool, best of luck!
First, a lot of the AL papers use _simulation_ scenarios rather than production scenarios, i.e. there is already more training data available, it just gets withheld. Obviously, if you already have more, you have spent annotating it, too, so there can't have been any saving.
Second, you always want to annotate more data than you have as long as the learning curve isn't flat, so it's not about how quickly you get up, but it's about should you keep annotating or does a flattening learning curve suggest you have reached the area of diminishing returns.
There are many sampling strategies like balance exploration & exploitation, expected model change, expected error reduction, exponentiated gradient exploration, uncertainty sampling, query by committee, querying from diverse subspaces/partitions, variance reduction, conformal predictors or mismatch-first farthest-traversal, and there isn't a theory to pick the best one given what you know (I've mostly heard people play with uncertainty sampling or query by committee in academia, but nobody in industry I know has told me they use AL).
Active learning becomes really useful when you hit diminishing returns, as most real-world ML applications deal with long tail distributions, and random sampling doesn't pick out edge cases for labeling very well. An easy way to tell if you're encountering diminishing returns is to do an ablation study where you train the same model against different subsets of your train set, evaluate them against each other on the same test set, and plot out the curve of model performance vs dataset size to see if you're starting to plateau. Or just eyeball your model errors and try to see if there's any patterns of edge cases it fails on.
Lastly, I'm pretty skeptical of model-based uncertainty sampling. In industry, almost every active learning implementation is very "what data should we label next," since model-based active learning is pretty hard to set up and confidence sampling is often not very reliable. That being said, I've anecdotally heard of some teams getting great performance from Bayesian methods once you have a large enough base dataset.
Anyway, here's my shameless plug for a post we wrote on the topic: https://medium.com/aquarium-learning/you-should-try-active-l...
Also, Aquarium Learning is just awesome. Super slick.
Can you shed some light on what you think are the most valuable methods for identifying high entropy examples for the model to learn faster? I'm familiar with Pool-Based Sampling, Stream-Based Selective Sampling, Membership Query Synthesis[1], but less certain which techniques are most useful in NLP.
Entropy selection for pool based methods looks at the output probability for each prediction of the model in the unlabelled data-set. Then it calculates the entropy of the distributions. (in classification this is a bit like looking for the most uniform predictive distributions) and prioritises those.
Entropy based active learning works ok but doesnt distinguish uncertainty that comes from a lack of knowledge (epistemic uncertainty) from noise. Techniques like Bayesian Active Learning by disagreement can do better. :)
Ima gonna hit u up too :)
Can you point to any big breakthroughs that have helped in recent years? Linear Hypermodels (https://arxiv.org/abs/2006.07464) seem promising, but that original experience has left me with some healthy skepticism.
In terms of breakthroughs in recent years, some things I'd point to would BALD (https://arxiv.org/abs/1112.5745) and its applications in deep learning as well. There has also been progress in coreset methods (https://openreview.net/forum?id=H1aIuk-RW).
I think that you're right that it used to be much to hard to get active learning to work. Part of what we're trying to do is make it easy enough that its worth the benefits.
I read the site, and no one likes cleaning data (no one), and answering questions from a "toddler machine" (for lack of a better term) doesn't sound as bad, but I was curious what potential trade offs there might be.
I tried HL, the experience was stellar (well done!) and it made me think...
To get AL working with a great user experience you need quite a bit of compute. How are you thinking about your margins, e.g the cost to produce what you’re offering versus what customers will pay for it?
this stuff helps us a lot on the margins point!
https://www.lighttag.io/blog/active-learning-optimization-is...
To be fair to the Humanloop folks, the criticisms probably don’t apply to the kinds of models their using (modern transformers).
I may update the post. thanks!
A slightly deeper intro: https://www.slideshare.net/nrubens/active-learning-in-recomm...
p.s. am the author of the above presentations; great to see Active Learning (AL) to finally get proper attention (I've been working in the AL area for 10+ years).
But for situations where there is a bigger pool of unlabelled data, active learning can identify which subset should be labelled to produce the best model performance, as long as the unlabelled pool contains "valuable" examples, the curve should remain steep and ideally meet performance targets much faster than for just annotating data randomly.
Also, there is some evidence that adding too many easy points to the training pool can reduce performance, see e.g. focal loss. So this could potentially mitigate that effect (mind you so could using focal loss, but that would require more labeling)
For some domains, with privacy concerns or rarity of objects, getting labelled data for deep learning is challenging.
There is decent research on sim2real i.e transferring models trained on synthetic data to real world applications https://arxiv.org/pdf/1703.06907.pdf
Synthetic data is particularly valuable when even the unlabelled data is expensive to obtain. For example if you want to train a driverless car, you may never see an ambulance driving at night in the rain even if you drive for thousands of miles. In that case, being able to synthesise data makes a lot of sense and we have lots of tools for computer graphics that make this easy.
Synthetic data can also be useful to share data when there are privacy concerns but my own feeling here is that there are better approaches to privacy preservation, like federated learning and learning via randomised response (https://arxiv.org/pdf/2001.04942.pdf).
In general though, outside of some vision applications, I'm pretty sceptical of synthetic data. For synthetic data to work well, you need a really good class conditional generator. E.g "generate a tweet that is a negative sentiment review" but if you have a sufficiently good model to do this, then you can probably use that model to solve your classification task anyway.
For most settings, I think synthetic data will work for data augmentation as a regulariser but will not be a substitute for all labelled data.
For the labelled data, active learning should still help.
At the end of the day, even in vision applications, real data is always better than synthetic data if you can get it. Things like sensor noise or interference are hard to replicate in synthetic data. Most teams turn to synthetic data for simulation purposes or as a last resort.
Some examples that we've actually worked on/are working on: * Contract classification
* Content moderation
* NER
* Customer review understanding
* support ticket routing
Can you talk about the tradeoffs or relationship between active learning and weak supervision from your point of view?
Short version:
1) we can get to the same level of accuracy with around 10% of the data points. Getting (and managing) a big enough data set to train a supervised learning model is the biggest thing slowing down ML deployments.
2) the model contacts a human when it can't label a data point with a high degree of confidence. You'll never have people with a bunch of specialist knowledge being asked to perform mundane data labelling tasks
Raza here, author of the post.
My high-level answer is weak-labelling overcomes cold starts and active learning helps with the last mile.
More detail:
We see weak learning as very complementary to active learning. By using labelling functions, you can quickly overcome the cold start problem and also better leverage external resources like knowledge bases.
But most of the work in training ML systems often comes in getting the last few percentage points of performance. Going from good to good enough. This is where active learning can really shine because it guides you as to what data you really need to move model performance.
At Humanloop, we've started with active learning tools but are also doing a lot of work on weak labelling.
Another consideration on the weak supervision side (for the snorkel style approach of labelling functions) is that creating labelling functions can be a relatively technical task, which may not be well suited for non-technical domain expert annotators for providing feedback to the model