What interests Reddit? A network analysis of 84M comments by 200K users
markallenthornton.com
markallenthornton.com
My pipeline right now looks like:
list of most popular SFW subreddits of all time -> gnu parallel -> curl on comment listings of top articles within subreddit -> json processing with jq -> gzip -> store on filesystem
So at the end I should have a directory of 5000 subreddits containing a gzipped json file of the comment threads of their top 500 posts.
I'm new to this kind of data processing but this is the easiest way I have found to do a short term, one time scrape of this data.
The exception is /r/all, which had infinite pagination, but the low API limits prevent much usability.
For example, this is what my get request to the REST endpoint looks like without URL parameters:
GET reddit.com/r/explainlikeimfive/comments/2w1xzo.json
Which returns the same response as appending .json to the the URL:
https://www.reddit.com/r/explainlikeimfive/comments/2w1xzo/e...
I have the raw data but it's infeasable to distribute due to size and the original source got in trouble for making it easily accessible.
Though, technology-wise, it is one use-case where SVG beats pixel graphics, both in terms of usability and interface (whether it is custom D3.js or something graph-oriented as http://sigmajs.org/).
http://i.imgur.com/ZLmWHrq.png
responsiveness of the community in order of most to least
- oldschool hacker (c/c++,bash,perl,regex)
- web dev (jquery, javascript, html, css)
- app dev (ios, objectivec, android, java)
Link "ate" the semicolon.
Also it took a lot of resources to calculate and we just didn't have the time to build efficient map/reduce jobs to do it regularly. It was done by hand in Mathmatica.
Google's PageRank is an example of an incestuous positive feedback loop. Web pages with many inbound links get ranked higher, and pages that rank higher get more inbound links (because that's how many people find things they link to).
The methods used there are awesome! Also, open source with a great website.
Seriously though, it's interesting how interconnected some things can be in this view. I'm not sure what sense I can make of those interconnections, though. Mousing around, while being a very frequent redditor (so my own neural network is making connections based on experience), I can kinda infer order out of things like the "government->state" topics connected to "force" and "property" among others (hints at the libertarian-leaning general population), and the "women" topic connecting to a whole host of stuff...the cyan colored section off to the top right might even kinda hint at the casual misogyny thing (which was a "ha ha only serious" kind of joke), with words like "bullshit", "logic", "proof", "assumption", "reasoning", and "evidence", being connected to "women" but not to "men".
But, without having spent years on reddit, and without my particular flavor of reddit (the subs I'm subscribed to), maybe I'd interpret the data very differently. I never quite know how to interpret network graphs like this, honestly, short of for things that are networks. i.e. a computer network topology on a graph shows useful data...the hops from one machine to the next. When connecting up one word to the next, it seems difficult to draw meaningful conclusions. Like my interpretation of the meaning of "government->state->property" as being a hint at the libertarian leanings of many subreddits, or the connection of "women->reasoning->evidence" as being a hint of many redditors belief that women are illogical liars (which is the impression many of my female friends have of reddit, in general, particularly when topics like date rape or the "friend zone" come up). Is that actually the context in which these connections are made? I wouldn't really know how to check. It'd be cool to be able to drill down to conversations in which the connections where made, but presenting that in a coherent UI seems challenging.
Perhaps the data was tailored when it was provided to the analyst, or it was censored after reception, but this felt too "PG-13" for an analysis of reddit's "interests".
Since 2012, Reddit operates as an independent company (Advanced Publications, the parent company of Condé Nast is a majority share holder though).
See: http://www.redditblog.com/2013/08/reddit-myth-busters_6.html...
Reddit is an independent entity, not a subsidiary of Condé Nast (like it used to be) and not a subsidiary of Advanced Publications (like it used to be).
It is an independent corporation, with it's own board of directors, and control of its own finances.
Just being a majority stakeholder doesn't mean you control the company either. There are a lot of details like share types and company by-laws that determine that.
And the grandparent saying that Advance Publications is "only" a majority shareholder is a little deceptive. The shareholders are Advance Publications, current and former employees (as part of a ESOP) and a small residual ownership of angels in the original company.
While it is true that a majority owner can't just do whatever it wants, the rules protect the financial interests of minority shareholders, mostly in the context of takeovers, not the editorial independence of employees. If Si and Donald decided they really didn't like the NSFW part of reddit I think they could get rid of it.
As for the ESOP percentage, all I've found is a reference in Forbes that describes it as a "sizable minority".
We don't even know how much influence AP has, since owning a majority share doesn't mean they have a lot of control. Also Reddit, Inc. just went through a round of investments, so those investors likely have a lot of influence too.
They were spun off as an independent company, but Advance Publications still has a large amount of equity in Reddit. I think it's because Reddit wanted to try a bunch of things that were too risky for AP.
A subsidiary is a company whose majority stakeholder is another company, which is exactly what reddit is. If reddit weren't a distinct corporate entity, they'd be a division.
Now, since AP doesn't have 100%, reddit isn't a wholly-owned subsidiary, but a company doesn't have to be wholly-owned to be a subsidiary at all.
It's still something to keep in mind.
I wish there was a publicly available dataset, but given all the privacy implications, I doubt it.
In fact I feel that a better way to see what redditors are interested in would be to just find (there may even be stats on reddit on this) the ~50 most active subreddits.
Maybe nothing could paint that picture, but if there are themes prevalent independent of the subreddit topics themselves, this kind of analysis could shed light on them.