Instagram’s Explore Recommender System
instagram-engineering.com
instagram-engineering.com
It definitely does. The homepage goes back to recommending a mix of my subscriptions and content generally related to them.
I like to reset my Youtube viewing history a couple times a month to see which direction my viewing habits will take my recommendations this time. I'll often ratchet into new territory. This month it was Warcraft 3: Reforged gameplay (Grubby), a game I haven't played in 15+ years.
The month before it was fiction book review channels. It's fun to change it up.
In Youtube's recent homepage update this month, they also rolled out a "Don't recommend this channel anymore" button on the dropdown which is very welcome.
This has been my experience as well. I end up using the "Don't show images like this for this hashtag" feature but they always seem to come back.
But yeah it can be off-putting.
I wonder if it's because if I see a thumbnail of something unwanted, I first have to open the post in order to access the menu, so the the app counts it as a view first, adding the unwanted post to things I "engage" with.
That aside, my personal experience would be vastly improved by having a "hide all images with text in it" option.
Yet the most critical input data - direct feedback from the user on the traits that shape their interests, is completely ignored/not collected.
Compared that to Spotify, whose goal I presume is to get me to listen to more music and buy tickets and merch through their occasional marketing.
I'm a music snob but damn does Spotify get me great recommendations on new releases, my discover weekly, and more. Not only that, I've bought tickets through their frequent listener promotions probably more than 10 times at this point.
I've been pigeonholed way beyond what I thought possible. Do other users really engage with the same 10 songs over and over and over that this is the default behavior of their recommendation engine?
I get "Discover" tracks which are from the same album I have downloaded to my phone!
I need a better diverse and robust recommendation system from Spotify (at this point I'm addicted, I listen maybe 5 hours on average) and a lot of times I get the same n number of songs again and again.
99% of the weeks Discover Weekly has come out, they have 1 or 2 real nice songs but the rest are the same "garbage" I've been listening to for a while.
Anyways, I agree, Discover Weekly needs a revamp.
The same thing applies for New Releases.
Also, I'm curious about the tradeoffs of revealing this information - does knowing this make it easier to game the instagram algorithm? From this article I'd think that having a more narrowly targeted account (for example someone putting selfies on one account and landscape photos on another) might make their embedding more similar to others. Another thought is that maybe someone liking a bunch of things unrelated to their content would make them wrongly appear in certain explore pages.
[1] https://fivethirtyeight.com/features/dissecting-trumps-most-...
EDIT: Maybe I'm just thrown off by the "Powered by AI" part of the article title. I was expecting more I suppose.
"look Ma i'm writing my own AI algorithm!!!" ... "writes linear regression by hand in python"
A. System should, without prompting, identify areas of improvement and innovation
B. Automated collection of data and the processing thereof, combined with application towards a concrete goal — does not qualify under A.
C. Part of A. is willing and unwilling discovery and exposure to both benevolent and adversarial environments and operating conditions
Bengio has a paper in 2003 that describes almost the same idea as word2vec (CBOW model to be exact).
There’s really no point in having a semantic argument, but just know that if you wish to do so you are arguing against many decades of wide usage of the term.
* scaling to more engineers/products: IGQL is an interesting way to compose ML pipelines with straightforward syntax
* scaling to more ranking candidates: an active user with a large follow graph who loads the explore tab likely has millions of eligible candidates - how do you load those fast? the idea of using a "distilled" model as a first, light ranking before using a full model as the final predictor is a good intuitive idea that I haven't seen described before.
* scaling KNN is hard: FB has done interesting work to make approximate nearest neighbor search fast, and opensourced it (the FAISS library which is referred to in the post). the improvements here are certainly non-trivial.
* scaling to more users: creating useful general purpose user embeddings is hard!
* scaling to more objectives: instagram has many business objectives, e.g. likes, follows, minimizing hides, so there is a need to have multiple models making many predictions. There is also a need to weight them intelligently, which is where the Bayesian optimization libraries come in.
in some sense, nothing is truly AI, but this is useful work which you can learn a lot from.
What makes this mildly interesting is the IGQL, but again, without knowing the full syntax, it feels pretty restrictive.
Youtube is doing much more advanced stuff, like Reinforcement learning@Scale, as comparing to Instagram in this regards.
This makes sense w/ what I see in IG recs: past behavior is strongly reinforced w/ littler diversity. Filter Bubble/Pigeon Hole problem.
So in conclusion, I would argue that the IG explore tab doesn't have ANY explore at all!
Seems like they are moving towards a structured RL implementation. There are elements of it, a follow-up post on some components would be interesting.