Network visualization of 50k blogs and links
graph.henryn.ca
graph.henryn.ca
[1] https://jgaa.info/accepted/2015/NocajOrtmannBrandes2015.19.2...
[1] https://pytorch-geometric.readthedocs.io/en/latest/generated... [2] https://umap-learn.readthedocs.io/en/latest/
Your visualisation tool may require it in a specific format, but it still about properties of your data.
You can see clusters forming of websites that talk about similar topics, like crypto, rationality, Canada, India, and even postgres!
The visualization was made entirely in webgl with some neat optimizations to render that many lines and circles.
The nice thing about a personal project is that I can do whatever I like with no constraints, so I built one that's suited for this project and fits my tastes.
[1] https://cambridge-intelligence.com/how-to-fix-hairballs/
I put up many links posts, so I probably link to an abnormally large number of sites.
The current visualization only shows the current state of the crawl, so it won't know about all of the posts.
Nice analysis! However, I'm guessing these arent your fav blogs as there are tens of thousands of entries! How did you decide which blogs to index, did you use some central registry of blogs?
One nice feature that would be helpful is the ability to preview the blog.
My mentor at the time had a traceroute dataset of the Internet and wanted to render it on top of Google Maps. I implemented a MapReduce algorithm that geolocated the data points and then produced Google Maps tiles at various zoom levels to show how the Internet was connected. It was pretty cool to visualize how the data flowed throughout the world and to be able to "dig deeper" by zooming into the mess of connections. Very similar to what this project does!
The project didn't go anywhere but it was a cool fun experiment and a great learning opportunity for me (S2 geometry is... well, weird, but touching MapReduce and Bigtable were invaluable exercises for my later tenure at the company). Those were very different times. I don't think you would be able to pursue such a "useless" project as an intern at Google these days.
Dataset was something from CAIDA, like this: https://www.caida.org/catalog/datasets/ipv4_prefix_probing_d...
IIRC we used the LGL algorithm (https://lgl.sourceforge.net/) while pinning any nodes we could get geolocations for, giving a nice hybrid geo/topo layout
I don't remember exactly how we got the geolocations, but often network routers have 3-letter airport codes in their DNS names, so maybe that? We may also have had a lookup table in el googz somewhere
Definitely a project whose time should again come! ;)
Thanks for chiming in. Good times!
Have you thought of a front end that is basically just text/plain HTML (in normal size) + navigation links to explore the blogs in one frame, and the currently chosen blog in another frame? That way, you could look at the blogs while travelling your crawl graph, a kind of "blog explorer".
And I have my own internal links visualization, which might be a bit over the top (GPU recommended): https://taoofmac.com/static/graph
Yep this is only for stuff that we've crawled, so we can't detect all of your links. Because we have limited crawling resources, we rate-limit the crawling by domain so we don't get stuck in spider traps. The current visualization only shows the current state of the crawl, so it won't know about all of the posts.
adithyabalaji.com
My question for you is how can I see what sites link to me, as opposed to what sites I link to?