Snorkel AI: Putting Data First in ML Development
snorkel.ai
snorkel.ai
With compose, a user defines a labeling function, and then compose scans the historical data looking for training examples to train a machine learning model.
The library has evolved as we apply it to more and more real world use cases, but it is based on approach in this paper from 2016: https://dai.lids.mit.edu/wp-content/uploads/2016/08/07796929....
The main lessons learned from this exercise helped us identify where our efforts would be shifted when using Snorkel. Of course there's never any free lunch, but Snorkel has what I believe to be a very reasonable and effective trade off. Snorkel provides both a huge decrease in overall costs, but critically it shifts costs towards the front of the development process. Writing a good set of labeling functions is a non-trivial piece of work. It requires the data scientist to have deep domain experience or a few fairly large blocks of time collaborating with and learning from a business user that is a already a domain expert. It has the upside though of forcing the data scientist to get a solid foundation of this domain knowledge, which I feel often times is underestimated in many ML projects.
Anyways, congrats to the team! Looking forward to checking out your future work.
My core problem is a multi-label problem, but my snorkel data, from the LabelModel is inherently single-label (mutually exclusive). What is the prevailing recommendation to do multi-label w/ Snorkel? Is the below what you are currently recommending?
For a given, k-wise multi-label problem:
1. Generate k binary datasets w/ LabelModel 2. Train k separate binary classifiers for each respective dataset 3. At inference/prediction time pass input though the k classifiers and get scores.
Is this what the current recommendation is? Create a set of binary classifiers?
Is it just a pipeline system with some helpers for running a couple of ML related functions?
Is it UI based?
Where do you run it?
I know these areas well and got nearly nothing from reading the splash page on the site.
1. Generate many different nosiy labels for your data by writing functions. These don't need to be correct, but they should make uncorrelated errors. They are basically domain knowledge you have of your data.
2. Snorkel takes the output of these functions, and based on their (dis)agreement, builds a generative probabilistic model to uncorrelate your labels, which may have had some overlap in the errors.
3. You train your final discriminative model on the output of that probabilistic model.
So, the main idea is to create many noisy labels instead of relying on a single high-quality label and Snorkel does the hard work of figuring out how to smartly combine these labels so you can train on something clean.
Part of the high level description, though, is that a lot of different parts and lines of work are integrated into Snorkel Flow beyond just this original programmatic labeling idea. So also programmatic operators for data augmentation, "slicing" or partitioning of data, and the overall end-to-end platform (UI + SDK) supporting iterative development of ML models via this paradigm of programmatic training data.
Edit: and a bit of fuzzy rule systems. Which just goes to suggest that I am probably well out of my depth.
When you look at ML models as commodities and the fact that you spend most of your time getting data, cleaning data or labeling data it leads to what they call Data Programming. I imagine this will be a UI where you can manage your dataset, by monitoring something they call Critical Slices.
The way he broke it down was either you can incorporate rules into your data or into your model. Because we want our model to be as general purpose as possible, it turns out you can squeeze some extra performance by "bronze/copper" quality data with handwritten rules in your dataset.
You can think of the model getting an extra boost from the latent knowledge within the rules.
Snorkel itself has been a open source package for a while - https://github.com/snorkel-team/snorkel
This new announcement is about Snorkel Flow
Imho, Snorkel kinds of tools ("weak supervision") are game changers for ML .. though the biggies get all the press. So I'm excited to see this end to end direction taken by the team.
The down-stream discriminates model's goal is to generalize via supervision.
I find in practice, especially w/ user-interaction, system most problems are not single-label, but multi-label. Also, in the single-label setting it's often necessary to define an negative "OTEHR" class which is very difficult to define w/ snorkel in my experience.
I signed up for a demo of the new platform, looking forward to chatting. Me and a colleague from work spoke w/ Henry last year about a potential partnership but I guess it got lost in the mix...
https://arxiv.org/abs/1703.00854
I can't say I followed all the proofs, but it seems that under certain limited assumptions about labelling functions they prove their generative function can do well.
Reading Snorkel it initially sounded like magic in the bad way, but this does make it clear that if your labelling functions are garbage or have certain kinds of problems there's nothing they can do about it.
Even leaving aside the generative model I think the focus on function-based data bootstrapping is great, which is why I've been following Snorkel's projects for a while.
It was extremely complicated to get it to do anything beyond the demos, and I was never successful to get it to do anything useful.
I ended up implementing some of the ideas myself, but I can't say I had any great success.
That said, I don't see anything here that would prevent you from using a pre-trained conv net as a labeling function, but I expect that multiple conv nets trained on a small corpus of data would be biased and make correlated errors, which violate their assumptions.
This looks super powerful in some cases, but I'm just not seeing how it can possibly generalize to every ML problem.
- As you imply, a lot has to do with the available sources of input signal- whether labeling functions, or 'transformation functions' for data augmentation, or other ops we've worked on... the input is obviously key.
- For data modalities like image, video, etc: Often the most successful approach is to (A) rely on some pre-processed features or "primitives" and write labeling functions over these- as my co-founder Paroma in particular has published about over the years- and/or (B) use metadata
- External models are definitely expressable as labeling functions, and we've worked on exactly that problem of modeling (local) biases and correlations!
For example, lets assume you want to identify something like a lung tumor. So you have many MRI images and they're all largely the same template of image. Using traditional image processing software like open CV, it's suprisingly easy to do more coarse grained tasks programmatically, like say, search this image for any circle that's brighter than the surrounding tissue and has a radius greater than say x.yz mm. If you find one, that function returns True if not False. That x.yz mm number is what you get from your radiologists that you work with to help you develop the labeling functions and this is just _one_ of the labeling functions. But basically it turns out if you construct a few of these functions with the help of domain experts and then use those functions all together with the information theory research the Snorkel folks do, you get pretty damn good performance!
That's btw why a lot of examples of ML today are ones where data is (i) simple for non-experts to label, (ii) non-private and therefore easy to outsource for labeling, and (iii) low rate of change (e.g. images for self-driving, basic NLP stuff for chat bots, etc)- this kind of data can be labeled cheaply and once, so hand-labeled training sets are (barely) economically feasible to build manually. However, most data is not that easy or cheap to label, needs to be relabeled constantly to adapt to change, and thus the investment in a programmatic approach is often far better even if certainly not push-button!
One important and practical answer that we've found: with an approach like in Snorkel Flow, you can inspect the source of the training data and correct it if biased- which you just can't do with e.g. a million hand labeled training data points. So in practice this is a big advantage we've found.
On the theory / research side, this is definitely an area we want to pursue further!
I'm imagining a dumb example like recipes, where "1 tsp salt" is a common format for ingredient. I'd imagine that the majority of ingredients follow that format, so it'd be a natural function to write. I'd also imagine that there's a correlation between following that format and being a recipe with a european background.
Generalize that a little bit, and almost by definition the simplest N rules that get the most coverage will cover the majority cases best. Being outside the majority cases is probably correlated with most "human issues," defined however you want. Being an artifact of the properties of what the simplest N rules cover, I'm not clear whether it'd be defined as local or systemic in the sense you've worked on.
I'm curious whether this falls under the theory you've worked on already or the theory you're talking about pursuing in the future. If it's something you've worked on already, I'd be very interested in reading what you have.
We've actually done some recent work on this (https://papers.nips.cc/paper/9137-slice-based-learning-a-pro...) where we have users define these critical "slices" approximately so that the model being trained can pay special attention to them (extra representation layers) so they don't get drowned out by the majority subsets/slices. But definitely a lot more to do in this area!
[1] https://github.com/FeatureLabs/compose [2] https://github.com/FeatureLabs/featuretools
- https://www.youtube.com/watch?v=yu15Nf5eJEE (14 min)
- https://www.cs.ucla.edu/upcoming-events/cs-201-jon-postel-di... (1hr talk - snorkel's predecessor was deepdive)
I went down the rabbit hole in this space about 3-4 years ago, and really got the message around "Dark Data". he's onto something huge and I regret not pursuing it further due to self doubt. hedge funds should be eating this up as well.
- Where to find more about the core Snorkel concepts: We've published 36+ peer-reviewed papers, along with blog posts, talks, office hours, etc over the years (see https://www.snorkel.ai/technology and https://www.snorkel.ai/case-studies), so I'll defer somewhat to those... but of course, academic papers can be painful to read (even when you wrote them!), so happy to also answer questions here.
- What Snorkel Flow is: Snorkel Flow is an end-to-end ML development platform based around the core idea that training data is the most important (and often ignored) part of ML systems today, and that you can label, build, and manage it programmatically with the right supporting techniques. This is based on our research at Stanford, where we spent several years exploring the basic question: can we enable subject matter expert users to train ML models with things like rules, heuristics, and other noisy sources of signal, expressed as "labeling functions" and other types of programmatic ops (ex: 'label this document X if it overlaps with dictionary Y'), instead of having to hand-label training data. This type of input, often termed "weak supervision", ends up requiring a lot of work to deal with as it is much noisier than hand-labeled data (eg the labeling functions can be inaccurate, differ in coverage and expertise, have tangled correlations, etc) but can be very powerful if you model it right! And Snorkel Flow specifically is focused on actually making the broader end-to-end process of building and managing ML with programmatic training data usable in production, rather than just on exploring the algorithmic and theoretical ideas as was the goal of our research/OSS code over the years!
- Why train a model if you have a programmatic way to label the data: In Snorkel, the basic idea is to label some portion of the data with labeling functions (usually it's hard to label all of the data- hence the need for ML), and then use ML to generalize beyond the LFs. In this sense Snorkel is an attempt to bridge rules-based approaches (high precision but low recall) and stats learning-based approaches (good at generalizing). This is also useful in "cross-modal" cases where you can write LFs over one feature set not available at inference time, but use them to train a model that does work on the servable/inference time features (e.g. text to image is one recent example https://www.cell.com/patterns/fulltext/S2666-3899(20)30019-2). But, of course, we believe in an empirical process all the way, which is another reason we like the Snorkel approach: if you can write a perfect set of labeling functions, then great- you don't need a fancy ML model, stop there!
- Does Snorkel work??: As an ML systems researcher, I'm always a bit perplexed by this question... the relevant questions for any system or approach are usually 'When/where might it be expected to be useful, and what are the relevant tradeoffs?' We've done our best to answer these questions over the years with theory, empirical studies, etc (see links above), and of course its very case specific. But one thing I'll note is that Snorkel is not a push-button automagic approach that takes in garbage and produces gold. It's our attempt to define a new input / development paradigm for ML--one which we've shown can often be orders of magnitude more efficient--but like any development process, it requires effort and infrastructure to use most successfully! Which is a big part of why we've built Snorkel Flow- to support and accelerate this new kind of ML development process.
- Who uses Snorkel? A few that have a published record: Google, Intel, Microsoft, Grubhub, Chegg, IBM... and many others at very large and smaller orgs that are not public
- What is going to happen with the OSS: The OSS project will remain up and open under Apache 2.0, same as all of the other research work we've put out over the years! See our community spectrum chat for more.
It's very powerful.
the labels are noisy, but you have a quantity that you could not get by humans, AND at a faster/cheaper rate. they provide analysis arguing that, for discriminative models, quantity CAN outweigh quality.
to your point it's not typically used w/ the image-only modality. It's mostly used where there is some meta-data attached.