Visualizing relationships between Python packages
kozikow.com
kozikow.com
You might be alyso interested to know that PyPI download stats are also available in BigQuery:
https://mail.python.org/pipermail/distutils-sig/2016-May/028...
(project recently unveiled by Donald Stufft)
I also added your article to the compilation at https://medium.com/@hoffa/github-on-bigquery-analyze-all-the..., thanks for your prolific series of analysis!
You can take a look at the table at https://bigquery.cloud.google.com/table/the-psf:pypi.downloa... and check out tabs "Schema" and "Preview". Mix of country_code, details* and tls* could give you "rough" grouping, but there would probably be 100s of users in the same bucket.
It is probably designed in a way to not let you target single user that way. It's likely that there is a pipeline scrubbing small buckets to avoid privacy leaks. I had to write pipelines like this myself on different datasets.
Also, PMI shows up in many contexts, including word2vec. Your measure is at best hacky (no direct interpretation, leaves many small nodes without links).
Otherwise, a nice work! But it would be great if visualized more interactively and aesthetically (yeah, I can be biased here).
- Mouse over to highlight neighbours
- Zoom in and out
- Browser based
- Ideally "hide this cluster" functionality
I wanted to add some of those features using d3, but the weekend ran out. Also "mouse over" would require me to move the vis from canvas to svg.
But: there should be exports from igraph to D3.js, and from Gephi to http://sigmajs.org/ (<- this is interactive, and higher level that D3.js; I haven't used it, though; with mouseover and zoom).
My another problem was that due the graph size it took ~1 minute to stabilize[1] on my browser and it was very wobbly and janky. I had to pre-load nodes positions and disable animation to make it usable for viewers. The fully automated analsysis of big graphs would require some server side component pre-computing the graph.
1. It was stabilizing quicker with smaller collision (newly introduced in d3 v4 force layout) parameter, but nodes were more likely to overlap.
Some rough idea would be high correlation of their neighbour weights, but low direct edge, but that doesn't work in some corner cases. More in https://kozikow.com/2016/07/10/visualizing-relationships-bet... .
"a small, fluffy rissun run on a tree" -> "rissun" is something like a squirrel.
If you want to have something more straightforward, look at adjacency matrix squared ("friends of friends") or some other measure of two nodes having a lot of common neighbours.
In a week I am having more time (and tutoring for 2 weeks a students on word2vec/glove visualizations), so I would be happy to talk about it.
It was hard to balance number of nodes well:
- Too few: Interesting clusters like robotics, numpy and openstack start to disappear, as they get dominated by massive clusters like django
- Too many: It's hard to see things.
What's missing is the ability for a user to "separate" clusters into spaces that they care about. Ideally though, the graphic should be able to display the relationships in a way that surfaces visually at a glance.