Why You Should Not Build a Recommendation Engine
datacommunitydc.org
datacommunitydc.org
I also don't like the idea it takes mysterious, scary "hefty data science" to be able to do recommendations.
- If you're recommending based on one thing (i.e. "people who viewed this also viewed.."), you'll be doing cosine similarity on the vector of viewers (i.e. the columns of the user x item matrix).
- If you're recommending based on many things (i.e. "recommended for you"), you'll be doing a matrix factorization of the user x item matrix. Pick SVD or NMF, depending on how sparse the data is.
- You probably won't be doing content based recommendations, without doing breakthrough machine learning research. For a lot of content, no-one really knows how to do this.
Better to analyze the latent factors, and then use that analysis to gain engineering-type insights into the structure of the problem, and target specifically what's driving people to like something. Exactly as Netflix has been doing in recent years. Exactly as people have done when writing books or developing other entertainment content for centuries.
Just dumping unsupervised clustering on the end users is a poor technique that should stay in the 1990's where it belongs.
I tried a number of off-the-shelf recommendation systems, none of which worked for the comparatively dense matrix we have (since we only have a few hundred products). I finally rolled one in house which took a few days of work (thanks, redis!) and has a bunch of custom filters that OTS systems wouldn't ever have.
I hate building stuff like this when OTS is available, but everything was overkill, or too hard to configure, or initially gave bad results and required too much black-box tweaking. It's been a success and was way easier than I thought.
Manually picking stuff is time consuming and might become stale quickly while a basic recommendation engine will take you a day or two. The whole point is to simply expose your consumer to more of your other products in the hope that they engage with one. If you have a wide scope of products, presenting the same tiny number of manual picks would be insane. Editors picks only works if you have a tiny number of products.
So build one if you want. And don't be afraid to make a slightly shit one either. As it's often better than tying up your time manually picking stuff out when you could be doing something better with your time.
The feature bit is correct, but the rating bit might not be. People very often give high ratings to things they are not interested in. People may highly rate jewelry that they cannot afford, or give a five star rating to a movies considered classics, that they don't actually want to view.
What you are trying to optimize is ultimately purchasing, continued subscription, ad views, or some other profit metric. That can be hard to directly measure, so something like ratings can be a stand in, but you may accidentally optimize for something deeply undesirable like recommending expensive jewelry to everyone that they will never buy.
You'd need to define
- a similarity measure for users (based on, e.g., what they bought and what they looked at) and
- you just recommend the top items the k nearest users bought (weighted by distance), possibly subtracted by the items already bought by the user in question and cleaned from explicit items.
Of course, the system would start really bad but it would get gradually better. This doesn't sound too complicated nor too simplistic nor "heavy data science".
Unfortunately there are no real technical insights in the article justifying why I shouldn't do this, so anyone can tell me why my approach would be bad or good?
What can be done with good applied math for making recommendations has variety nearly beyond belief; the approaches outlined by the OP are, of course, just silly.
For the cold start problem, that need not be very difficult -- just start in niches.
value = k1 * users + k2 * products + k3 * products * users
Prioritize building your rec engine feature after you've built user on-boarding features and product on-boarding features because without users and products, rec engines contribute 0 value.