Deep Neural Networks for YouTube Recommendations
research.google.com
research.google.com
There's an interesting presentation of how it's created on SlideShare
http://www.slideshare.net/MrChrisJohnson/from-idea-to-execut...
The addition of Discover Weekly really confused me. Shouldn't the features that create a radio station from an artist or a playlist fill this need already? Why is it only updated weekly? I haven't tried other services much but it feels like Spotify isn't doing as much as they can with recommendations.
I wonder if there's any data on how common this is. I listen to large shared playlists or the radio feature the vast majority of the time to try to find new music.
Either way, radio-from-artist or radio-from-song can't utilize your listening history. It's quite possible you listen to an artist for different reasons than the reasons most people listen to that artist - in that case, you will get recommendations based on the majority's reasons.
You would expect personalized recommendations to have potential to do much better, and I'd argue that's exactly what we see with DW.
After glancing through the top few, I quickly went for the search box. To their credit, once I do a few searches, the recommendations drastically improved when visiting next time.
A side effect of this, of course, is that you can study all kinds of stereotyping and biases by repeating my experiment in various regions I suppose.
Yes, I watched a bunch of Dota replays during a recent tournament. No, I don't normally go on youtube. So I watched a daily show video clip that was linked. All my "watch next" and "recommended" are Dota. That's not smart, that's aggravating. I would have been ok watching a couple more ds clips, but instead exclusively bad recommendations were made based on poor data.
I dislike the idea that my world gets filtered by algorithms, but I really hate when they're obviously bad at it. Although I suppose I should be grateful that it's easily spotted when it's bad?
n conjugation with other product areas across Google, YouTube has undergone a fundamental paradigm shift to- wards using deep learning as a general-purpose solution for nearly all learning problems.
Can you talk about how this works in practice? Is the deep learning group separate from other teams and then tackles problems from different areas as needed, or are there deep learning engineers in each project area that are building nets for each different area? Is the ML team also redesigning product architecture by building products around reinforcement learning?
A recent article [1] revealed how engineers are trained in ML across Google.
[1] https://backchannel.com/how-google-is-remaking-itself-as-a-m...
Features about the videos such as titles and tags, as well as features derived from audio and video, are introduced in the ranking phase.
Does this mean something different from feeding the age of the video, relative to when the training example was recorded? Feeding in the age of the video seems like a fairly obvious idea and like it should train the network to favor newer videos. If it actually means how long ago the training example was recorded that is rather strange, as I don't see how that would be needed on top of the video age. Neat graph, there.
I am often annoyed at how overly focused online recommendations systems are for my overly specific recent trends, rather than broader interests I display over months or years of using a product (looking at you Amazon). It seems like it should be relatively easy to learn 'this guy likes little video essays about art and science and sometimes fun talk shows' and yet YouTube has been pretty bad at recommending such video-essay style content to me. Perhaps this will improve it, although I wonder how much the recent history features end up overwhelming overall years-long type data about what interests me broadly and not just yesterday.
As an aside, is it really "Deep Neural Networks for YouTube Recommendations" if you are using 5-ish layers of embedding, ReLu units, and output? A bit humorous, that.
I tend to think the focus on recent behavior is an artifact of underfitting. Research into richer temporal modeling is needed and recurrent networks seem promising.
We debated internally whether to use the "deep" moniker - Alexnet was 8 layers, so maybe the threshold is 8? The depth seems sort of irrelevant since stacking layers is trivial once the basic architecture is in place.
That is an information bubble. The algorithm cannot detect low quality or populism, neither can it recommend opposite standpoints, and at the end of the day it has a real effect on a country's politics and the well-being of many people.
Do you have means of quantifying such effects? What are possible countermeasures?
If you cannot talk about that, then this would be my feedback: Perhaps you could train a language model to find opposing views in video titles and tags and then diversify the video recommendations based on that.
'Information bubbles' have existed as long as people have had a choice of newspapers to buy and TV channels to watch. Calling for Youtube to artificially 'balance' videos seems like political interference.
If you think relativism is fine. Then "As opposed to planting your flag in the ground that your camp is always right and the outgroup is evil?" is also fine.
See how silly and immediately self-contradictory relativism is?
With recommendation engines, your bubble, without effort, ossifies.
google does right to not chose between left and right, moral or imoral.
For example that no death threats are spoken when a daughter defies her farther's will who she is supposed to spend time with. That caricaturists, satirists and atheists are safe. These are things that western cultures have established, and which could arguably be endangered by letting in refugees by the millions and by prohibiting cultural criticism at the same time. I am myself not convinced of the urgency of this threat, but I think these are some of the more convincing arguments against Merkel's refugee politics. Other arguments are for example second order effects or equilibrium effects, e.g. that conservatives, professionals and business folk amount to a counter-reaction that is worse than letting in refugees in a more controlled way (i.e. Brexit and brain drain).
> exactly the same people who hate gays and feminism in the first place.
I have no idea about the numbers, but I am pretty certain both groups exist. Those who use these arguments as pretense and those who are honestly concerned about the efficiency, safety and trust our culture has established (which e.g. allow us to focus on education, art and science).
The system suggests me lots of click-baits and low content quality videos (with massive views though). It's very rare that i get a great video that i eventually really enjoy in my recommandations.
My guess is that the system can't really tell if the video itself is made of good quality, brings good and fresh content. Is that the case? how do you guys rate and measure the intrinsic video quality?
I wish the recommendation engine had a better idea of what I liked based on the fact that I've been using Youtube for years, and I've thumbs-upped a lot of videos, and told it a lot of channels and videos that I don't like. But maybe that's just asking too much?
So, each video is mapped to fixed size vector of floats? A user's history is now a matrix of size [number of videos, embedding size]? What are the other parameters in this sentence "Importantly, the embeddings are learned jointly with all other model parameters through normal gradient descent back propagation updates."? And how do you concatenate all these into a "wide layer" when users would have histories of different length?
This is of course not optimal, as the network should be able to learn how best to summarize the sequence. In the paper, however, we emphasize the importance of withholding certain sequential information from the classifier.
The recommendation system can't seem to handle outliers but maybe that's asking too much of current technology.
- the recommendations are very often not interesting to me because
+ they cater to the lowest common denominator (you won't believe these 10 hilarious fails, PewDiePie picks his nose, etc.)
+ I have already watched the video
+ a video has been in the Recommended section for weeks and I haven't clicked on it. What makes you think I'll change my mind after several weeks? If I don't click a video within a couple of days of it appearing in the section, it's a dud. Don't keep showing it
+ the video is from a channel I am already subscribed to. That's not a recommendation, it's trivial and not helpful
+ most or all of the videos in the section are sometimes matching the same key word. I once clicked on an Amy Schumer video, and for many days every video in the Recommended section was a Schumer video. This is terrible. The same thing happened after I clicked a Craig Ferguson video.
- the feedback UI is not streamlined. I have to click through multiple menus to be able to say: not interested in this channel
- there should be list of key words that I can specify where if the video matches one of them, don't add it to the section. Conversely, there should be a list of key words that when I specify them, the recommendation engine goes out and looks for videos matching them, and then adds some of them to the section
I love watching interesting and creative how-to videos (DiResta, Tested, etc.), but even after several years of watching them, the recommendation engine seems to not have caught on to that.Is the deep learning approach already deployed for regular users? I have not seen a change in the quality of the recommendations.
Sorry to sound so negative, but I think this is a huge wasted opportunity. There is tons of amazing content on youtube, and it's often very hard to find.
I am very interest in this. Deep NN are quite an interesting subject and something I'm personally quite curious about.
I also use youtube recommendations quite a lot for some fairly specialize interests [which I'll keep unstated for now]. My current impression has been that the recommendation system has only gotten worse in the last ten years and is now nearly broken (I get recommendations from third party websites now).
As I recall things, Youtube removed most user recommendation controls 5-10 years ago and the guesses it makes still haven't made up for this loss.
But there are other things I find even harder to understand. I find that when I'm not logged in, after choosing 5-10 videos, youtube will start to recommend good stuff, indeed things that I'd like on my regular recommendation list but which I never do see there.
My impression of my regular recommendations is that serves nothing but crudes averages, videos that I just assume someone pays Google to recommend. ("Sports" "celebrity fails", etc).
Which brings me to shock that the cream of the cream of AI somehow deploys this to me. I get that Convnets have made quantum leaps in image recognition competitions. AlphaGo was a clear advance. But where is the progress here? If the recommendation engine is categorizing videos, either the categorizations don't correspond to my experiences or its using the categorizations incorrectly. Broadly, my impression is the algorithm is swayed by whether a video is broadly popular rather than whether its in a given category. And I work hard to prune every off-topic suggested video or suggested topic, yet I get what seems like poor to worthless quality recommendations.
Please make it so I can block specific YouTube users from EVER being recommended to me, or showing up in sidebars or whatever.
I feel these features might actually start being useful to me if there was a way to tell YT "Please, please stop showing me this users videos, I absolutely never want to watch them at all".
My personal curation > your algorithms.
How do you decide what the N in Top N should be?
I see you guys scaled features yourselves, why not use BatchNorm?
Do you think you could have eliminated the manual feature engineering with some learnable feature engineering? I'm mostly thinking of some sort of parametric activation functions, but I'm curious if you've thought about it.
Any thoughts on the Wide & Deep paper, did you try incorporating similar ideas?
Did you experiment with LSTMs for turning watch/search histories into fixed vectors?
You guys trained a regression model, whereas the common wisdom is that neural nets aren't so hot at regression, did you try training this as a bucketized classification problem?
Again, thanks for the paper and taking time to answer questions :)
I'd say recommendations have gotten somewhat better, if a little too clickbaity still (you saw one video with a squirrel? here are ten squirrel video compilations!).
Second question: Did humans at some point assign names to DNN-established clusters or vector elements or what they're called? Sometimes I get OK recommendations, but with a really bad label (for instance a 100% minecraft LPer recommended as an example of a "shooter").
Thanks for making your hard work available. It is very interesting from a technical point of view. I'm struck by just how huge a challenge this is given the enormous corpus size.
Are you familiar with Joe Edelman's work?[0] He specifically uses YouTube recommendations as an example of many algorithms designed to use the wrong metrics which leads to undesirable outcomes for users.
Have you ever looked into attributing reasons to users' visits? It seems likely that many users aren't looking for general recommendations that blend their entire use of YouTube together but want specialized recommendations linked to why they visited YouTube this time.
[0] https://medium.com/@edelwax/is-anything-worth-maximizing-d11...
I like recommendations about music, and I like recommendations about non-music stuff, but they're all mixed togather and that's painful.
It's a vicious clickbaity cycle.
Methods that work better for a population as a whole might not work better for a large subset of that population, and might even cause users to stop using features entirely. The lack of transparency in recommendation algorithms combined with the homogenizing effect of distributing low-quality content this way is something I find somewhat depressing.
But perhaps there is a more risky strategy that takes longer to craft and actually delivers hours and hours of content to the user (but needs to fail longer before getting there).
Does anyone know whether RL is used for recommendation in practical settings, and if so what is the current state of the art?
We have struggled with interpretability, both while debugging mistakes made by the system and exposing plausible "reasons" to users.
There was a fascinating discussion [1] about interpretability during a deep learning panel at KDD this year.
If I was a huge fan of books, movies, music, youtube picks of another user, it may be there is a deeper connection of the kind of quality we are both looking for, and so his or her recommendations would be highly relevant.
I think the issue of a system determining whether you like the steampunk genre vs the quality of only that particular steampunk video is separate from the issue of ratings.
I am not happy with YT recommendations because they suggest crap videos to me and not the finest one available for that topic, just as he said.
The system should rather suggest me a different topic but with the best quality/content available, rather than a super similar video with crappier quality/content.
Which usually just leads to curation systems being key. And they work well, until they are gamed. And they will be gamed.
Heh. Me too. I suffer from joint/ligament issues and buy supplements for those.Now, Amazon thinks I am a geriatric and recommends me 'helping hand' sticks and incontinence products :\
Reminds me o this tweet
https://twitter.com/kibblesmith/status/724817086309142529
Amazon is a $250 billion dollar company that reacts to you buying a vacuum by going THIS GUY LOVES BUYING VACUUMS HERE ARE SOME MORE VACUUMSThat's what makes the problem so interesting! Most recommendation systems are terrible, and those that aren't, are good only for the first 1-2 recommendations.
And then there's Google Search, which so thoroughly demolished existing search result recommendation systems (remember Altavista?) that they now own the market and are one of the most valuable companies in history.
When you finally solve a recommendation system problem in a way that actually works, it's a huge freaking deal!
[1] http://www.slideshare.net/xamat/recommender-systems-machine-...
People were literally bizarred by youtube, saying they were there by recommendation. (I have this video in the recommendations also...)
Sometimes YouTube recommends me videos with clickbaity gross thumbnails or from YouTubers I dislike or have no desire to watch but there is nothing I can do to to stop it recommending these to me, why can't I just go to these users profiles and block them and have them removed from my YT experience?
Block just seems to stop people from messaging you, not from you being shown their videos by an algorithm.
If you used to have a bad recommendation system, and then you switch over to this system, then it will still be trained with data generated by users who saw the old recommendations, leading it to have a bias towards the same bad predictions.
Is there any way around that?