This is a fantastic post. Would love to see an even more detailed walk through some code examples, and discussion of development to production of these models. Do you have any other resources you would suggest?
But to answer you question here: For production ANN we have Annoy integrated into our backend as a service. Annoy was an easy choice bc it checked out box for JVM support.
For training the models we have endless amounts of behavioral data, so we didn't even need to look at transfer learning. For this query2vec example, it was trained on 1 year of queries which takes 15min/epoch on an AWS p2 GPU. We do all our preprocessing (heavy normalization) in pyspark.