402 karma · joined March 19, 2007
What I'm looking for and what this post doesn't cover is moving beyond the binary case. Looking for Multi Positive and Unlabeled Learning...
I signed up for a demo of the new platform, looking forward to chatting. Me and a colleague from work spoke w/ Henry last year about a potential partnership but I guess it got lost in the mix...
the labels are noisy, but you have a quantity that you could not get by humans, AND at a faster/cheaper rate. they provide analysis arguing that, for discriminative models, quantity CAN outweigh quality.
to your point it's not typically used w/ the image-only modality. It's mostly used where there is some meta-data attached.
My core problem is a multi-label problem, but my snorkel data, from the LabelModel is inherently single-label (mutually exclusive). What is the prevailing recommendation to do multi-label w/ Snorkel? Is the below what you are currently recommending?
For a given, k-wise multi-label problem:
1. Generate k binary datasets w/ LabelModel 2. Train k separate binary classifiers for each respective dataset 3. At inference/prediction time pass input though the k classifiers and get scores.
Is this what the current recommendation is? Create a set of binary classifiers?
The down-stream discriminates model's goal is to generalize via supervision.
When you look at ML models as commodities and the fact that you spend most of your time getting data, cleaning data or labeling data it leads to what they call Data Programming. I imagine this will be a UI where you can manage your dataset, by monitoring something they call Critical Slices.
It's very powerful.
I find in practice, especially w/ user-interaction, system most problems are not single-label, but multi-label. Also, in the single-label setting it's often necessary to define an negative "OTEHR" class which is very difficult to define w/ snorkel in my experience.
This makes sense w/ what I see in IG recs: past behavior is strongly reinforced w/ littler diversity. Filter Bubble/Pigeon Hole problem.
So in conclusion, I would argue that the IG explore tab doesn't have ANY explore at all!
Jobs are scheduled (Azkaban) for reruns/re-training and pushed from data env to the feature/model-store in live env (Cassandra). Online models are exported to SaveModel format and can be loaded on any TF platform, eg java backends.
Online inference using TF Serving. Clients query models via grpc.
A lot of our models are NN embedding lookups, we use Annoy for indexing those.