Help Reddit build a recommender
reddit.com
reddit.com
This was originally the "hard problem" at the center of Reddit.
Let me explain what I mean by that. There used to be a quaint notion that to be a respectable tech startup, you had to have a "hard problem" (technologically speaking) at your core, which you had an innovative "secret sauce" solution for, preferably one you were patenting. After all, if not, then someone can just copy you and squash you like a bug, right?
Since then, YC's insistent focus on making something people want, Eric Ries' lean startup gospel and many entrepreneurs' own experiences have thankfully gone a long way to convince people (most importantly SV investors) that focusing on a "hard problem" is not only unnecessary, but may end up being a fatal distraction.
This is a pretty good example of how the "hard problem" can turn out to be completely irrelevant. Once it was clear that the recommendation engine wasn't a growth vector, the Reddit team seemed to drop it out of sheer pragmatism. They just needed to keep the site running.
I can't recall many who cared or even noticed that the "recommended" tab was gone. But from that point on, Reddit was more free to become not just a quirky "personalized news" startup, but what it has aspired to since: the front page of the internet. And only now, just now, do a good chunk of the millions of users think a recommender might be nice.
It's the startup version of "you aren't gonna need it" — if it doesn't drive growth, push it aside.
You're correct that reddit didn't/doesn't need any form of recommendation engine for individual items -- that's what voting + subreddits are -- it most definitely did (and still does) need something for recommending subreddits. This is a problem they should have been addressing from day 1, it isn't make or break but it takes the site way beyond its current usefulness to hardcore users.
Hardly a day 1 problem.
Hopefully it will be something people want, because I want it as well. But if it is not, I'd rather he pivoted into something where he can succeed than taking his startup to the ground because of me. A startup that fails helps nobody.
InfoQ taped it, but they may not release it for up to 6 months.
And yes, it was Bradford Cross.
I suspect the "hard problem" idea might not actually be that antiquated. Most of the newer batch of startups are either not profitable at all or not profitable enough to justify VC investment. There's almost no technology risk, but the tradeoff in market risk is such that even if you succeed, you're not as profitable. This is the trap Reddit is in.
I've used methods known as collaborative filtration, whose goal was to estimate how a given user would rate a given item basing on knowledge of preferences of other users of similar interests. The initial scope included a naïve Bayesian classifier and a technique called Slope One [1]. The latter one is particularly interesting as according to claims of its authors allows to make a very good estimation in a very short time using solely a very simple linear model. The preprocessing is both time- and space-wise expensive though as it requires you to build a matrix of deviations between rated items.
After reducing the data set to a single subreddit and filtering it from users who weren't avid voters I ran the algorithms and after some tuning I was very content to see promising ROC curves and decent AUC values. Models built around NBC and S1 achieved comparable results when it came to such metrics as precision, recall and F-measure.
When I went to discuss the results with the professor teaching the class I've heard "That's indeed promising, but how about comparing those results with a really naïve model which would just take an average of existing votes by a given user?". Guess what: the model built solely using a single call to the avg function was nearly as good as the NBC and S1 models.
Now I understand why the guys from Reddit are looking for external help with the recommender. It's a way less obvious task than it might seem to be.
[1] http://lemire.me/fr/documents/publications/lemiremaclachlan_...
Edit: s/machine learning/data mining/
1) Say you estimate (as you propose) that a user will always give their average rating. This might get you good-ish error and ROC as a prediction task, but will give zero recommendation value because the prediction for a given user will be constant for all possible recommendations.
2) Say you estimate that a user will give the average score that the item has received across all users. Again, possibly good-ish in terms of prediction ROC and RMS error, but this offers no personalization (all users get the same predictions, i.e. you're basically just showing the default Reddit ranking).
Both of these baselines are vastly inferior to even really stupid models like "how many times have I upvoted stories from this submitter" in terms of recommendation value, but the latter is (if I recall from my own experiments) much worse when evaluated on the basis of overall ROC.
I would strongly suspect that a correctly implemented NB or S1 would vastly outperform either of the two baselines in terms of actual recommendation utility (even though when you look at the baseline's ability to predict actual numbers, they might be comparably good in an RMS sense).
The moral of the story: one must be very careful when trying to quantify the performance of learning systems; actual utility is often difficult to evaluate merely by looking at standard statistical measures of accuracy.
Turns out increasing the dimensionality of the input 17 thousand times just reduces the amount training data for each attribute. Duh :)
Gremlin (https://github.com/tinkerpop/gremlin/wiki) works great for real-time recommendations.
See "A Graph-Based Movie Recommender Engine" by Gremlin's creator, Marko Rodriguez (http://markorodriguez.com/2011/09/22/a-graph-based-movie-rec...)
(Content models are the other, probably less interesting, 50% of the solution.)
Reddit's old interface doesn't work anymore now that there are so many subs. The fact that there has been so little interface improvements in the last couple of years is pretty sad. I can't imagine browsing the site without RES.
The way things work now only helps to magnify the lower quality trend because the homepage gives undue weight to content from popular subs.
I think there could be some interesting UI solutions to this problem. If more people treated the Reddit API like the Twitter API, there could be applications that aren't necessarily supposed to replace the traditional Reddit browsing experience, but to make whole new experiences (e.g. Flipboard).
How to you measure success? After I create my algorithm how do I know that I'm close to what reddit wants? Without answers to these questions, IMO, this is an exercise in futility. I'm not close enough to the project but written my fair share of classifiers and clustering engines any machine learning problem there needs to be a way to measure success. My point of view on a great result is different from reddits for sure.
If a story had tags and there was a system where the frequency of tags appearing in a subreddit mattered it would allow me to look at r/startups and then find the other subreddits relevant to my interests.
reddit made the mistake of treating every subreddit as its own individual isolated community without considering crossovers in interests. If tagging existed then this would not have been a problem. Today 6 years on it's still impossible to find good subreddits relevant to specific interests, tags would have been one of the solutions for that.
Check it out: http://ec2-50-16-106-77.compute-1.amazonaws.com/ - (it's still in its infancy)
You can link the same post as a reply to multiple items. This allows for complete flexibility. Posts that are relevant to more than one section can live in each of those places
I think that this was one of the most important strategic decisions in reddit's history, and that they got it right.
I'm not saying tags can never work, just that any proposed tags system needs to supplement, not destroy, the siloing of subreddit communities. And be simple to use, even for the 99% of redditors who never even vote or subscribe to anything.
I'll happily admit that tagging could not have grown Reddit to anywhere near its current size without the site collapsing on itself. The different feel to each community (compare F7U12 are AskScience) is much more appealing than a single homogeneous group. However, I think that it was the first step towards breaking the promise to create an personalized news aggregator. I for one was disappointed when the recommended posts feature was dropped.
The initial missteps with whitelabel sites like the Wired-branded reddit and lipstick.com are amusing in hindsight. I'm not sure how reddit with a pink background with Courier as the primary font was supposed to attract a female audience.
For example, if a post could be tagged "startups" and was posted to r/business, when I tried to find other subreddits besides r/startups about startups I could search "subreddits with x or more stories tagged "startups"" and I'd be presented with r/business.
They wouldn't exist as a replacement for subreddits, they'd exist along side and serve as a way to connect subreddits by topic. Subreddits currently exist as their own entities with no crossover which doesn't work well for expanding a users subscriptions to other subreddits relevant to their interest.
That said don't listen to me. I have quit using reddit, except for /r/gonewild.
[1]: http://youtube.com/
For one thing, many good stories languish on the "new" page and never get enough votes to get a fair shake. Collaborative filtering doesn't help with this, if anything it makes it worse.
Last night I made a crude boomerang by glueing two rulers together, this morning it had set and my son pressured me to try throwing it before I'd even finished my breakfast. Right when it started to curve, it hit a telephone pole and broke at the glue joint.
When I see many of the things people want to do on reddit, my first impression is it will wind up like that. For instance, LSI is one of those things that does not work so well in real life... They still seem to be teaching kids about it, but not that you get results almost good doing dimensional reduction with a random basis set.
If you've got some semantic analysis and predictive models, you can make an automated system that picks quality relevant content out of the "new" queue and because you can use smart feature selection you don't need to wrangle as much data -- training is orders of magnitude faster and you don't need to futz around with hadoop.
Then the similarity of r/programming to r/coding would be based on two numbers:
b = number of people subscribed to both r/coding and r/programming
n = number of people subscribed to r/coding
similarity = b/nYou could take it a step further and incorporate more than explicit up/down vote features, such as "clicked", "commented", "saved", etc.
Then incorporate some business rules that filter recommendations by subreddits, boost results by time, and now you have a decent recommender.
Easier said than done of course.