Movie Recommendations with k-Nearest Neighbors and Cosine Similarity
gist.neo4j.org
gist.neo4j.org
I've recently done a similar research on finding "company peers" using company descriptions. The descriptions are put through the LDA (http://en.wikipedia.org/wiki/Latent_Dirichlet_allocation) to find the topics expressed in each description. Then, similar companies are identified using K-L divergence (http://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_diverg...) between their topic distributions. It worked much better than cosine similarity in my tests.
My guess is that it would work better for movies too. If interested, have a look at my results here: http://akuz.me/2014/03/finding-company-peers-using-lda/
Look in to hill climbing algorithms if you're ever curious how basic optimization algorithms (re: games) work.
http://radimrehurek.com/2013/12/performance-shootout-of-near...
Most of the challenge is in getting a way of assessing the value of innovations in the algorithms - how do you know how well it works ? Difficult unless you are running a large scale recommender that users can't opt out of (because pop goes your stats if the do!)
You could say, "I want to watch a movie like The Wolf of Wall Street" and it would find the closest 10 movies in the graph.
It's still something I'd like to play with if I find the time.
Most of my time thus far has been spent gathering the dataset, but I do have a few example cypher queries answering the following simple questions [2]:
1) What actors have appeared in the most AFI Top 100 films?
2) What are the genres of the top ten films?
3) Have any actors appeared in 2 or more of the top 25 films?
I'm working on building a much larger data set using a combination of freebase and imdb so that I can have enough data to start exploring much more interesting interesting questions (e.g. graph the frequencies of genres over the past 60 years; for a given film, find movies with the greatest overlap in genres, actors, and directors; generalize the n-degrees-to-bacon problem to work on any two actors; etc).
[1]https://github.com/mcphilip/film-graph
[2]http://htmlpreview.github.io/?https://github.com/mcphilip/fi...
Its not based on imdb, but based on http://www.gnovies.com
It wouldn't provide accurate prediction of the best pick movie to watch, but it might come up with an indirect, quality pick that might otherwise never been seen.
Which other relatively simple techniques could we apply to find out who is similar to us in terms of movie taste, etc?
Also, It may be better to compute cosine similarity based on movie type, as it may provide less noise.
Moreover, I have never used R but it seems like a very neat language...
For anyone who's studied a bit of ML its simple though.
http://nicolemargaretwhite.blogspot.com/2013/12/movie-recomm...
You can do it in Python but you're going to have to persist and query your data somehow, and I think at that point you'll "get" why a graph DB might be beneficial. It's not an ML-related thing though, it's a data query performance thing.