Is this pretty much a solved problem or is there more to explore?
Is this pretty much a solved problem or is there more to explore?
Currently, I would expect that teams are experimenting with repurposing Attention-like architectures in some way to get better embeddings, especially from sequence like features.
That query / item tower is cheap and cache-friendly. These tend to be used for high recall.
You can echew those performance benefits in favor of neural networks that allow for the features of the query and item to interact. These models are goaling for high precision.
I would consider two towers pretty much a necessity for large corpus retrieval as step 1 in any rec system with many items and many requests. Stage 2 or later models can be heavier and be whatever you like.
Evaluating recommendation systems is hard because you actually require a human in the loop. Even worse, giving the recommendations alters the human behavior. Then you need to think what metric are you going to use. For training you will most probably use a proxy metric that correlates. Maybe you want to optimize different metrics and they actually need to be balanced. Then there are lot of confounding variables: maybe a better UX will improve the metrics than a better algorithm, or a change of products.
It seems that with big enough data you can improve old models with deep learning but I think recommenders are very far from similar gains to other fields (NLP or CV for example). And most companies don't have that much data.
For smaller customer bases this is a tricky problem, but I’d argue that automated recommendations don’t work at a small scale anyway so manual curation is king.
Doing this automatically is a huge positive ROI thing over manual curation. In fact, humans are not even that good at coming up with good recommendations. Manual re-arrangement has almost universally been an anti-relevance feature in A/B tests I've looked at.
But of course, almost nobody does the automation right.
That said, this pytorch work looks to live more in the realm of application than academia. Building and scaling large neural embeddings is pretty close to industry practice these days and this library at least claims to solve some of the challenges in doing so.
https://paperswithcode.com/datasets?task=recommendation-syst...
Like many have already said, this is mostly an academic answer (even if some of the papers are written by the industry).
In the industrial world, the answer is a lot more subtle. Each domain has its own constraints and the best method will vary.
Also, keep in mind that in some domains like ad tech, the whole measurement process is messed up by the attribution mechanism, which attributes a sale to the last click (which clearly is a poor indication of whether the recommendation engine is doing a good job -- the industry settled on this attribution mechanism because it is easy to audit).
(Disclosure: I work at a a company which just expanded their product portfolio in this direction. The first few pilot customers show very good results, but there are a lot of aspects and nuances and alternative approaches we haven't had time to try yet.)
You will only see the field make steady, regular and infinite progress towards an unknown asymptote.
And so Netflix' recommendation has become largely irrelevant now that the company is pushing its own content agenda.
Their recommendation algorithm is operating under massive override by marketing rules.
Asking your friends for recommendations. In my opinion automated systems are still not nearly as good as another person knowing you and your preferences.