A Graph of Related Subreddits
anvaka.github.io
anvaka.github.io
Just wanted to say thank you for sharing this! Would be happy to answer any questions - graphs are my long time hobby, and I love them!
PS: You can find more recent graphs and fun projects here: https://twitter.com/search?q=from%3Aanvaka%20min_retweets%3A...
edit: never mind, I just read your GitHub readme. But the question still stands as if "users posting in x, also posted in y" is a good way to infer similarity. Could comparing top-ranked posts be a better comparator?
That said, there are a few subreddits that are too popular and similarity results were too saturated (/r/videos, /r/funny, etc.) so I did a manual override by looking into most commonly mentioned other subreddits, and sometimes into `about` blurb of subreddit).
Please don't consider these recommendation as source of truth! It's just a fun way to discover other subreddits :).
I'm also very open to change this metric to something else - please let me know if you have any recommendations!
[1]: https://github.com/anvaka/sayit#the-data - describes the data, indexing scripts are here: https://github.com/anvaka/sayit/tree/master/scripts
[2]: Manual overrides can be found here https://github.com/anvaka/sayit-data#sayit---recommendation-...
The great thing is that it’s actually more actionable as far as recommendations go! Everybody has already heard of the bigger version of this subreddit, but they probably haven’t heard of the smaller versions. And it’s self-correcting. As a subreddit gets bigger we are less likely to recommend it (which is great because it needs our help less)
If you guys are interested in seeing how your recommendation work for the entire reddit, I'd be happy to build you a spaceship similar to this one https://github.com/anvaka/word2vec-graph .
I couldn't find an easy way to download the entire recommendation graph, but it would be awesome if we could make it work. My email is the same as this account at gmail, and twitter is all open: https://twitter.com/anvaka
I fooled around a bit with lastfm data for band recommendation and found this sheet quite helpful.
If you are interested in learning more about asymmetrical similarity, here is a great primer by Tversky - http://www.cogsci.ucsd.edu/~coulson/203/tversky-features.pdf
My suggestion would be something much simpler ie. to try to compare content itself e.g. top 1000 posts from each subreddit, and estimate (say) cosine sim + Tfidf. Wouldn't that be a better indicator? Also, instead of pairwise comparison, you could try clustering (HDBSCAN for example) to reduce computational complexity.
But great work, love your visualizations!
I wrote on this topic [1]. My method [2] basically uses simple counts on edge weights, and then estimates the expected edge weight and its variance using Bayesian priors. It then attaches a t-score or p-value to each edge, and then you can filter out edges with too low t-score.
The idea is that weak edges can still be statistically significant if they connect "small" nodes. In any case, the library I wrote includes the implementation of a few other methods, in case they work better for your data type.
[1] https://arxiv.org/abs/1701.07336 [2] http://www.michelecoscia.com/?page_id=287
I think when I created this tool there was no recommendations on reddit.
When it was introduced later on reddit I was contemplating about using reddit's own recommendations, but at that time it was missing a few smaller subreddits, so I just put it of onto the shelf of projects to try.
I'm not sure of a fix for that, but would it be possible/helpful to weigh it by average upvote/downvote of the comments left from users of said sub? Meaning if sub A is about how much baseball sucks, and sub B is about how amazing baseball is, while determining if the 2 are similar you'd find out most posts from sub A to sub B are heavily downvoted and so probably not similar.
I'd still see ability to determine absolute value of relationship as a valuable property of a recommender
Is that really a good metric of similarity? Just myself, I post in several unrelated subreddits semi-regularly from programming to video games to music, art and even stone masonary, i've posted in subreddits for TV shows i've watched, or just completely random things.
I use reddit as a place where I can learn about and interact with people on nearly any subject or topic and I take advantage of that when I can. I'm sure my posting habits aren't that unusual. I'm just not sure that's really an accurate way to gauge similarity.
Jaccard similarity does not count only the number of people who posted to A and B, it checks how many people posted to A, how many people posted to B, and how many of those people posted TOGETHER to A and B. That togetherness gives us hints what is related, and after it is computed, we can divide by the total number of poster to both A and B (independently), which brings the value to something that we can use to compare against other subreddits. If that value is close to 1, it means that almost all users who posted to A have also posted to B. If it is close to 0, then the overlap is much smaller.
I share your interest in creating graphs (vide some older project of mine, a map of Stack Exchange tags https://p.migdal.pl/tagoverflow/).
When I dived into various visualization of Subreddits (as a background reading for this side-project https://observablehq.com/@stared/tree-of-reddit-sex-life), your vis was the best for exploration (out of many, many). The only other one that I found and this level was a Sigma.js-based https://www.jacobsilterra.com/subreddit_map/network/ (beautiful, but more static).
I see you are doing something with quantum tensors, which is very impressive! I'm still struggling with concept of regular tensors, and would love to have a good intuition/visualization for them - do you have any pointers?
I've found it illuminating to see tensor diagrams.
- https://medium.com/@pmigdal/in-the-topic-of-diagrams-i-did-w...
- https://www.math3ma.com/blog/matrices-as-tensor-network-diag...
Also, right now, we are developing a matrix visualization in https://github.com/Quantum-Game. However, it is pretty much work in progress; for a slightly more mature one, go to Quantum Game 2 website and in the element encyclopedia there is one.
I am always up for talking about tensors, so feel invited to drop me an email.
Animations are fun, but only if they convey artistic or intellectual meaning.
Also, the default zoom level is zoomed in too far (large text, not showing the wohle graph). perhaps this is because viewport resolution/size (phone vs desktop) is not taken into account when chooseing a zoom level (font size)
https://anvaka.github.io/sayit/?query=findasubreddit
Super interesting!
So many small and unknown subreddits, and one of the main connections is another subreddit called "Somebody Make This".
I'll take a peek under the hood later, but this on the surface is very cool.
Unrelated, but had to say it somewhere in this discussion: Thanks @anvaka for your open source graph libs! Way more performant than anything else I've tried.
Speaking of which, it'd be awesome if the site could automatically generate multi-reddits from the results. I think a multi-reddit constructed from that graph would be quite interesting ;)
r/chairs has 800 people subscribed. r/chairsunderwater has 115k. It's just reddit things ¯\_(ツ)_/¯
the algorith is non-deterministic, so if you follow the link more than once, you get a different spatial arrangement each time.
some are very pleasing, with identifiable clusters, and others are just a seemingly random scatter of names.
also interesting: https://anvaka.github.io/sayit/?query=lostredditors
(no lines are generated here; is this a feature or a bug?)
Edit: oh wow: https://anvaka.github.io/sayit/?query=Jung
As a side-note, I'm quite disappointed that r/DMT does not pop up when starting with r/JoeRogan...
The graph is extremely shocking and not at all what I would have expected.
A broader base of people (about 1/4 to 1/2 of the voting age US public) is aware and supportive of Donald Trump and for a wide variety of reasons, and much more than half at at least aware of Donald Trump. Jordan Peterson has a much smaller following, who are interested in him for similar reasons to each other.
or
"popping their filter bubbles", or "are generally interested in boundary-pushing ideas", or "are susceptible to the rage-inducing trolls who run fringe communities".
"Affinity for extreme ideas of any kind" maybe a stronger / more common personality trait than "interested in one extreme point in the vector space of ideas".
https://subredditstats.com/subreddit-user-overlaps/longevity
vs
I always keep on looking for new interesting subreddits. This tool is the best I saw for this task so far.
This implementation is tailored to smaller graphs with sometimes long text boxes.
I'll make a note to extract it to a reusable component.