HNHacker News
TopNewBestAskShowJobs

eggie5

402 karma · joined March 19, 2007

eggie5.com
submissionscomments
eggie5··on Query2vec: Search query expansion with query embeddings
Although your post is orthogonal to what we present in this work, I think it still merits discussion.

The high-level concept of a product is important for any commerce company w/ an unbounded and unstructured product catalog. In this case, the concept of a Dish is a very valuable primitive for Caviar and GH to understand.

You propose an interesting technique for dish tagging:

1. train word2vec on menu item text 2. Generate menu item embeddings by aggregating respective word vectors 3. Curate dish (label) set 4. Generate embeddings for dish labels by aggregating word vectors 5. Tag menu items w/ their dish label by NN search.

Then once you have this high-level concept of dish you can drive all kinds of interesting product innovations:

* trending asian noodle dishes * tacos in san Diego * It also helps w/ the item sparsity problem in matrix factorization

Do I understand?

I'm going to try this on our next learning Friday.

We have taken a different route for dish tagging. First we have a graphical model tag some items w/ low coverage but high precision and then we feed that into a classifier to generalize and provider full coverage. It's like a generative model fed into a discriminative model.

eggie5··on Query2vec: Search query expansion with query embeddings
thanks, I think we are in agreeance.
eggie5··on Query2vec: Search query expansion with query embeddings
I agree:

In the IR community you will see image2image search implemented by passing an image through a headless CNN and then do ANN search on the embedding space. Then that can lead to discussion about cross-model hashing which you suggested. I think also you touched on LTR techniques w/ siamese networks. For example we can use a siamese network to learn the diff between two embeddings.

In the Recys community for collaborative filtering you can take interaction data between users and items and decompose that matrix into two parts using SVD which is called the user and item embeddings. Also in the recsys community you see people doing item2vec (what I did) for the item embeddings and then aggregating those for the user representation (embeddings).

Maybe you will grant me the novelty of this method in making the connection between queries and converted restaurants as a similarity metric and then applying it to the IR use case of search query expansion? Should the post be more clear that I am not claiming to have invented item2vec or ANN (approximate nearest neighboors)?

eggie5··on Query2vec: Search query expansion with query embeddings
Author here. Thanks for the comment. This is a tough business to be in, but very exciting a the same time.

I'm trying to effect change by making meaningful contributions to our search pipeline by applying novel techniques from information retrieval literature.

For example, with this query expansion technique, we can service search queries in small markets (not a lot of restaurants) more completely by showing something instead of nothing. Or in big markets, we can service obscure, tail queries by fining something similar. This helps increase the overall recall of the search system, making it better!

Don't loose faith in GH yet!

eggie5··on Query2vec: Search query expansion with query embeddings
Correct, the model has a fixed vocabulary set at train-time. However, we can benefit from the fact that the distribution of queries in most search systems follows the Zipfian distribution. That means we can capture the vast majority of queries that are ever issued in a fixed vocabulary. Daily retraining helps us pick up new queries outside of the vocabulary.
eggie5··on Query2vec: Search query expansion with query embeddings
Author here. Thanks, for the recommendation. I actually got inspiration for this idea from chatting w/ industry colleagues at SIGIR. Last SigIR in Paris, I spoke to some people from Apple, where they are using a technique like this for the app store search.
eggie5··on Query2vec: Search query expansion with query embeddings
Hi, I'm the author of the post. We actually just presented this work and more at PyData NYC. We share more implementation details. Here are the slides: https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodg...

But to answer you question here: For production ANN we have Annoy integrated into our backend as a service. Annoy was an easy choice bc it checked out box for JVM support.

For training the models we have endless amounts of behavioral data, so we didn't even need to look at transfer learning. For this query2vec example, it was trained on 1 year of queries which takes 15min/epoch on an AWS p2 GPU. We do all our preprocessing (heavy normalization) in pyspark.

eggie5··on Query2vec: Search query expansion with query embeddings
Author here: We hope to fix this w/ the techniques described in the post! The search engine will first collect a high-recall set of candidates which is then passed to a high-precision ranker. This should help you w/ your cuisine and dish queries misclassifications. It will also introduce semantic understanding into the query engine meaning if you type "French" it will not give you "French fries" but French cuisine.
eggie5··on A dumb reason computer vision apps aren’t working: Exif Orientation
CNNs are translation and scale invariant thanks mostly to the pooling operation. Good data augmentation (rotating images for example) would have build a model more robust to this effect.
eggie5··on Ask HN: Who is hiring? (August 2019)
There's loss of fpgas in space already. Space micro is one example
eggie5··on Ask HN: Has anyone ever been hired from “Who wants to be hired?” threads?
I have
eggie5··on Fast.ai – Deep Learning from the Foundations
no, I have never watched a course before and watched the first lecture of this last night w/ no issue.
eggie5··on Differences between the word2vec paper and its implementation
instead of thinking about what it is in practice: skip-gram negative sampling, I think it's much more intuitive to think about what it is in theory: extreme multi-class classification.

word2vec is a multi-class classification problem with a softmax output layer and cross-entropy loss. The novel part of word2vec, in my opinion, is two:

1. dataset (proximal input word & output word) generation from documents eg: skiagram, CBOW, etc 2. engineering speedup for softmax: Approximate Softmax eg Negative Sampling using NCE, hierarchal softmax, etc

If you just build word2vec w/o step 2, it's a easier to understand. Then when you get that working, add in the negative sampling speedup trick which isn't core the theoretical algorithm.

eggie5··on Uber S-1
can you share the two numbers you are using for comparison?

The Q4 numbers I see:

Uber Eats: $165 million Grub Hub: $205 million*

* https://investors.grubhub.com/investors/press-releases/press...

eggie5··on Open Sourcing Peloton, Uber’s Unified Resource Scheduler
there's pipeline.ai, airbnb said they'd open source theirs this year and also, TFX suite is getting there. Platforms are becoming popular
eggie5··on Machine Learning: Full-Text Search in JavaScript – Relevance Scoring (2015)
he generated query-document features. Now he just needs to collect relevance labels for the documents, then he can learn a ranker a la LTR.
eggie5··on Implementing a Neural Network from Scratch in Python
+1 partial of the loss w/ respect to the weight
eggie5··on Mattermost Raises $20M Series A Funding
Probably a good call. Slack trains recommender systems w/ all user data mixed together. Although w/ attempts to preserve privacy... See their paper from recsys '18 Vancouver.
eggie5··on Who is horse_js?
cluster horse_js and Tom Dale's tweets in an embedding space and you can confirm your hypothesis.
eggie5··on DOJ: Hackers broke into an SEC database and made millions from inside info
This isn't hard to believe if you've worked w/ the Edgar system!
eggie5··on What Kagglers Are Using for Text Classification
I've seen a rule of thumb that if the ratio of samples to words per sample is less than 1500 you prob don't have enough data for embeddings/cnn
eggie5··on What Kagglers Are Using for Text Classification
Why not just use a 1d conv over the sequence of embeddings?
eggie5··on Ask HN: What did you learn in 2018?
* factorization machines * word2vec * WARP loss for Matrix Factorization
eggie5··on A newly discovered tea plant is caffeine-free
Rooibos brother,

This is my favorite blend I randomly found outside of the Schönefeld Airport in Berlin and have been importing the states ever since:

https://shop.tgtea.com/Rooibush-Cream-Caramel-001312-100/561...

eggie5··on Amazon Personalize – Real-Time Personalization and Recommendation for Everyone
Plesently surprised w/ the choice of algorithms AWS provides. They skip classic MF and go straight to the good stuff:

* Item-Item CF (workhorse amazon original from 2003 w/ modern enhancements) * Deep Pooling Models (à la Covington. YouTube Deep Rec.)

Although I'm a bit perplexed as to the difference between Deep-FM Recipe and the FFNN Recipe as it describes the same thing...

https://docs.aws.amazon.com/personalize/latest/dg/working-wi...

eggie5··on EuclidesDB: a multi-model machine learning feature database
the pieces are:

* CNN (plenty of pre-trained * Approximate Nearest Neighbors Database (Annoy, etc) * webserver to host CNN and serve UI

It's popular to serve tensorflow models w/ tensorflow severing + kubernetes and that's what' I've done in the past.

eggie5··on EuclidesDB: a multi-model machine learning feature database
The successive convolutional layers in popular CNN architectures learn representations from very general (edges) to very specific (dog breeds) from the input to the last conv. layer respectively. Depending on how different your new domain is will dictate on what layer you will take the CNN representation from (early layers or later layers) and whether a generic ImageNet model will work at all.

Of course if your new domain is very different than the distribution of the original training data, it is a good technique to fine-tune the network a bit with your data.

eggie5··on EuclidesDB: a multi-model machine learning feature database
I gave a talk on the theory behind image to image search if anyone is interested. Image search is essentially what this backend well suited for and what the graphic on their home page uses:

http://www.eggie5.com/126-semantic-image-search-video

eggie5··on EuclidesDB: a multi-model machine learning feature database
yes, this is essentially the Clarifai product: use a CNN (without the softmax) as a feature extractor for images and then do ANN search on them in the embedding space.
eggie5··on EuclidesDB: a multi-model machine learning feature database
If you've every tried to deploy a deep learning based image to image search product, you will know the engineering challenges especially with the Approximate Nearest Neighbors infrastructure. This is a good progress in abstracting out that step!
← PreviousPage 3 of 10Next →