Is it just a pipeline system with some helpers for running a couple of ML related functions?
Is it UI based?
Where do you run it?
I know these areas well and got nearly nothing from reading the splash page on the site.
Is it just a pipeline system with some helpers for running a couple of ML related functions?
Is it UI based?
Where do you run it?
I know these areas well and got nearly nothing from reading the splash page on the site.
1. Generate many different nosiy labels for your data by writing functions. These don't need to be correct, but they should make uncorrelated errors. They are basically domain knowledge you have of your data.
2. Snorkel takes the output of these functions, and based on their (dis)agreement, builds a generative probabilistic model to uncorrelate your labels, which may have had some overlap in the errors.
3. You train your final discriminative model on the output of that probabilistic model.
So, the main idea is to create many noisy labels instead of relying on a single high-quality label and Snorkel does the hard work of figuring out how to smartly combine these labels so you can train on something clean.
Edit: and a bit of fuzzy rule systems. Which just goes to suggest that I am probably well out of my depth.
Part of the high level description, though, is that a lot of different parts and lines of work are integrated into Snorkel Flow beyond just this original programmatic labeling idea. So also programmatic operators for data augmentation, "slicing" or partitioning of data, and the overall end-to-end platform (UI + SDK) supporting iterative development of ML models via this paradigm of programmatic training data.
When you look at ML models as commodities and the fact that you spend most of your time getting data, cleaning data or labeling data it leads to what they call Data Programming. I imagine this will be a UI where you can manage your dataset, by monitoring something they call Critical Slices.