HNHacker News
TopNewBestAskShowJobs

eggie5

402 karma · joined March 19, 2007

eggie5.com
submissionscomments
eggie5··on S3 Strong Consistency
does this mean the TF S3 plugin will stop cluttering my logs w/ s3 retries?
eggie5··on Email a Dumpster Fire
you forgot prototype before jquery
eggie5··on Positive-unlabeled learning (2017)
Just did a deep dive of this a few days ago. Mostly looked at the Elkan paper, which this post called naive. I've been researching in the context of weak-supervision data from the Snorkel software package. It's very easy to create positive labels but very hard to make negative labels.

What I'm looking for and what this post doesn't cover is moving beyond the binary case. Looking for Multi Positive and Unlabeled Learning...

eggie5··on Snorkel AI: Putting Data First in ML Development
Thanks for the lead on PU Learning.

I signed up for a demo of the new platform, looking forward to chatting. Me and a colleague from work spoke w/ Henry last year about a potential partnership but I guess it got lost in the mix...

eggie5··on Snorkel AI: Putting Data First in ML Development
you have a dataset of images and you write code (labeling functions LF) to label the images. Snorkel handles the pipeline but more importantly corrects the conflicts/correlations between the LFs. The output is a supervised dataset w/ mutually exclusive labels a la softmax classification.

the labels are noisy, but you have a quantity that you could not get by humans, AND at a faster/cheaper rate. they provide analysis arguing that, for discriminative models, quantity CAN outweigh quality.

to your point it's not typically used w/ the image-only modality. It's mostly used where there is some meta-data attached.

eggie5··on Snorkel AI: Putting Data First in ML Development
Snorkel has a label mutual exclusion assumption right?

My core problem is a multi-label problem, but my snorkel data, from the LabelModel is inherently single-label (mutually exclusive). What is the prevailing recommendation to do multi-label w/ Snorkel? Is the below what you are currently recommending?

For a given, k-wise multi-label problem:

1. Generate k binary datasets w/ LabelModel 2. Train k separate binary classifiers for each respective dataset 3. At inference/prediction time pass input though the k classifiers and get scores.

Is this what the current recommendation is? Create a set of binary classifiers?

eggie5··on Snorkel AI: Putting Data First in ML Development
Thanks Alex. I'm sure you can relate that when you have a an unbounded input distribution (like w/ user-interaction systems), defining that other class w/ current snorkel is difficult/impossible.
eggie5··on Snorkel AI: Putting Data First in ML Development
yeah, you can do single-label w/ snorkel, but not multi-label. Multi-label snorkel would be the killer feature bc making the negatives (ie for a softmax) is very hard especially when you work w/ user-interaction systems with an unknown negative distribution.
eggie5··on Snorkel AI: Putting Data First in ML Development
can you do multi-label w/ Compose? Snorkel only supports single-label.
eggie5··on Snorkel AI: Putting Data First in ML Development
can you do multi-label w/ Compose? Snorkel only supports single-label.
eggie5··on Snorkel AI: Putting Data First in ML Development
can you educate me on how you do multi-label w/ snorkel? As far as I can understand that's one it's largest drawbacks.
eggie5··on Snorkel AI: Putting Data First in ML Development
Why do think that? It's a very powerful and useful technique. You can get supervised labels in a day (ostensibly for free) vs paying humans to do it and waiting...
eggie5··on Snorkel AI: Putting Data First in ML Development
That's too bad. Myself and colleagues have had good success using the current snorkel package.
eggie5··on Snorkel AI: Putting Data First in ML Development
A human will label data according to hand-rules or heuristics. What's the difference is a program labels data according to hand-rules or heuristics.

The down-stream discriminates model's goal is to generalize via supervision.

eggie5··on Snorkel AI: Putting Data First in ML Development
The former project is a python package that labels your data using a weak supervision technique. It's not just a pipeline, it's a sophisticated algorithm that helps combine multiple competing labeling functions by removing reweighing based on correlations vs a naive majority-voting scheme.

When you look at ML models as commodities and the fact that you spend most of your time getting data, cleaning data or labeling data it leads to what they call Data Programming. I imagine this will be a UI where you can manage your dataset, by monitoring something they call Critical Slices.

eggie5··on Snorkel AI: Putting Data First in ML Development
To that end the OS project will remain: https://spectrum.chat/snorkel/general/announcing-snorkel-flo...
eggie5··on Snorkel AI: Putting Data First in ML Development
Snorkel does weak supervision for you. It takes your unlabeled data and use defined labeling functions (LFs), maps the LFs on the data and then de-correlates relates everything to give you a dataset that you can use for multi-class single-label supervision.

It's very powerful.

eggie5··on Snorkel AI: Putting Data First in ML Development
I can say Grubhub and Chegg use it.
eggie5··on Snorkel AI: Putting Data First in ML Development
anyone figure out how to do multi-label classification w/ snorkel? It seems like it's current formulation only supports single-label, ie softmax.

I find in practice, especially w/ user-interaction, system most problems are not single-label, but multi-label. Also, in the single-label setting it's often necessary to define an negative "OTEHR" class which is very difficult to define w/ snorkel in my experience.

eggie5··on America is stuck at home, but food-delivery companies still struggle to profit
why does everyone say Grubhub is not profitable?
eggie5··on Shirt Without Stripes
simple LTR w/ clickstream data would fix this easy
eggie5··on TripAdvisor cuts hundreds of jobs after Google competition bites
this killed Trivago too
eggie5··on Booking.com agrees to EU demands to change travel offers
True the are lots of booking scrapers like trivago
eggie5··on Ask HN: How are you using Machine Learning?
good
eggie5··on Building a Search Engine from Scratch
You can use a knowledge graph for query expansion. For example, if the query is "Dan Dan Noodles", you can expand that to "Asain Noodles", "Chinese", "Tan Tan Noodles" to achieve higher Recall in your results.
eggie5··on Instagram’s Explore Recommender System
So IG ins't using collaborative filtering? The whole process starts w/ simple NN search in the account embedding space. Those candidates are then passed to the ranking stack.

This makes sense w/ what I see in IG recs: past behavior is strongly reinforced w/ littler diversity. Filter Bubble/Pigeon Hole problem.

So in conclusion, I would argue that the IG explore tab doesn't have ANY explore at all!

eggie5··on Understanding Serde
I recently saw this strange word Serde in some Hive table create commands...
eggie5··on Ask HN: What is your ML stack like?
Development of models in our data environment: notebooks, pyspark EMR clusters for analytical workloads and offline models, tensorflow/EC2 P2s for online models.

Jobs are scheduled (Azkaban) for reruns/re-training and pushed from data env to the feature/model-store in live env (Cassandra). Online models are exported to SaveModel format and can be loaded on any TF platform, eg java backends.

Online inference using TF Serving. Clients query models via grpc.

A lot of our models are NN embedding lookups, we use Annoy for indexing those.

eggie5··on When your data doesn’t fit in memory: the basic techniques
one word: Dask
eggie5··on Query2vec: Search query expansion with query embeddings
The latent graph structure is encoded into the embeddings. Each node in the graph is a vector and you traverse the graph by preforming NN search in the embedding space.
← PreviousPage 2 of 10Next →