1,004 karma · joined December 23, 2014
The later graphs use data that's already not public (what tags users visit together, and what countries tags are visited from), so there's no reason not to use visits instead.
The graphs for visits over time do look very similar (in general question traffic by tag roughly matches questions asked, but as a slightly lagging indicator)
Perhaps there's a way I could make that clearer in the post!
https://stackoverflow.blog/2017/05/16/exploring-state-mobile...
https://twitter.com/drob/status/875493967865008129
Perhaps it could play a role as a more indirect effect: people leaving Google/Microsoft/Apple/etc, starting new companies or joining senior positions at others, and spreading this particular practice.
This does not mean the effect isn't confounded with some other factor, but it does mean it's not a multiple hypothesis testing issue.
You can explore the code and regression yourself! Take a look: https://github.com/dgrtwo/tabs-spaces-post
Random forests are a method that's often effective in taking into account many interactions among high dimensional data.
During a hackathon it can be hard to tell when to keep searching for an easy solution like that, as opposed to going with something slow you know will work- sometimes it turns out to be a dead end.
Thanks for the recommendations!
1. That's the way we were thinking about it :)
2. Oh, excellent! We hadn't found that or we'd have used it, and we'll start working with it.
3. Tomorrow I'm going to blog about how we approached the machine learning. Short version; we manually came up with regular expressions to classify a training set based on titles. The idea is that when we experimented with manual annotations on titles, the vast majority of the time we were looking for only a few key words. There's no question that this adds biases and will not be entirely accurate, but manual inspection convinced us it was a good enough approach for our hackathon, and most of the articles we identified with the resulting algorithm would not have been found by the title regex alone.
You can see the table of regular expressions [here](https://github.com/dodger487/analyze_hn/blob/master/topics.c...) and a bunch of (pretty unstructured) analysis code [here](https://github.com/dodger487/analyze_hn/blob/master/hn-analy...).
I'm working with DataCamp to develop an R course that covers dplyr, tidyr, and other newer additions to the R language.
Our goal is to solve those problems, and it is not a zero-sum game.
But note that precisely because git/github is used in combination with almost all other tags, it doesn't have a high correlation with any particular technology. A correlation (roughly) means "If I know you use tag X, you're more likely to also use tag Y." But knowing someone uses git doesn't let you guess what other technologies you use, because as you note they can use almost anything.
In short, if you're connected to everything, that means you're correlated with nothing.