HN long term change
jacquesmattheij.com
jacquesmattheij.com
That said, I think an article about economics or politics can be more profound and deep than a lightweight article about technology and be better for the tone of the site. My personal preference is for news here to be technical, but an article being technical is not enough for it to be interesting. Overall, however, I want articles with meat; that dig a little deeper than "Head First SQL" or "How I made a blog engine with Erlang" (or "I am Phillip Greenspun and I don't like people in Northern California").
While I love beautifully designed languages as much as the next guy, I've seen probably 50 blog articles posted here about "Why language design matters" or "Why language design doesn't matter" or "Why there can never be a better Lisp". While these may be technical in nature, they're often very shallow and redundant.
As far as I know, it would be an impressive feat to determine depth automatically, but I think it would give you a better picture of how the tone is changing over time in a relevant way. And maybe it would be a good filter for submitted articles.
* Objectively ranked categories based on word usage, not author-provided tags
* Legible graphs with fewer colors
* Analysis that weights comments by karma (or, better, a reasonable non-linear function on karma)
* Open source for reproducibility and better outside critique
Edit: In fact, I'd like to see it enough that I might build it. Any other ideas?
Maybe determine "popularity" first by number of articles and then again by score.
I have to do all kinds of tricky date parsing to figure out when an item was posted so this is not as easy as it seems at first glance.
There seems to be some confusion about how to interpret the graph, the Y axis is 'per post', the x axis is per block of 1,000 posts.
The graph would have been harder to read at 138x1000 pixels so I've stretched it a bit.
When this is all done I'll make all the data available so that people can mine it for what it is worth.
A fair amount of work went in to this little project, not all of it automated (unfortunately), a lot of time went in to making sure the tags would be somewhat relevant.
If there would be any major trends that I could not explain by looking at samples of the data I would have certainly investigated.
I hope that if there are such trends that I missed that they will come out in the follow up (weighted by votes, see above), if not then the analysis will have to be much more detailed, and that will probably mean a lot more handwork than what went in to making this graph.
The 'good news' for me is that there is no unbounded growth of the 'unspecified' category, that would be a fairly large indicator of trouble.
The bigger issue is the fact that this is just everything that is posted and not flagged, so it is if you wish a view of the 'new' page, it has nothing to do with the 'home' page, I'll try to address that tomorrow.
As for the labelling and clustering, that was based on keywords in the title from a fair sized sample, and from the urls the links pointed to.
What I am specifically searching for is larger trends, smaller trends would be very difficult to catch using this method.
I'm actually quite surprised how even the graphs come out over the longer term, I would have expected more variation in the submissions.
So if there is a problem at this point in time I would conclude that the problem is not in the submissions, they seem to have roughly the same subjects over the long term as they did in the beginning, with the exception of a shift of focus away from 'startups' in the first year or so of operation.
I think that has to do with an influx of programmers / people interested in technology in general whereas originally most of the people on news.yc were active in the startup scene.
That's been my impression for a long time. Do your techniques allow you to measure the trend of people complaining about the site deteriorating? Because that's been going on for a long time too, and in approximately the same way (though possibly in cycles).
I think that has to do with an influx of programmers / people interested in technology in general whereas originally most of the people on news.yc were active in the startup scene.
Pretty clearly that is because the site was originally named Startup News and had a relatively narrow scope, then was renamed to Hacker News as part of explicitly broadening the scope.
No, especially not because plenty of those get flagged and die.
Any data on how other similar sites (e.g., reddit) have changed?
Even the cyclic nature that edw519 referred to in the 'new years exchange' seems to be mostly limited to the voting, it has nothing to do with the actual submissions.
This was probably part of the iterative change from 'Startup News' to 'Hacker News'. I can't recall exactly when the name changed, or whether it was a response to that trend or precipitated the wider focus.
Edit: Change was made 14 August 2007.
http://news.ycombinator.com/item?id=1047482
Meanwhile less topical, but more sensational stories shot to the top.
Also, what is "as khn" and "as kyc"? It looks like one replaced the other.
ask hn (ask hacker news) and ask yc (ask Y Combinator).
ask yc was obviously more popular in the early days.
Those two should be summed, but the misspelled on is used very rarely (fortunately).
I'll re-do the graph tomorrow when I'm awake, it won't affect any other rows or the shapes though.
There were many subcategories as well, but I've used only the top level of the tags to make the graph legible.
In total I used about 200 different tags.
But he also has lots of stuff that is not so easy to categorize, so that ended up depending on the ease with which the title let itself be identified either under 'technology' or, in the worst case under 'blogs'.
A similar problem appears with the 'mainstream' media websites, and it was solved in the same way with the top level category as a catch-all after other matches were ruled out.
Example: (Is Amazon EC2 oversubscribed)
As for the scale, it doesn't get much more precise than this, the only concession to legibility is to stretch the graph horizontally because otherwise it would be only 138 pixels wide, vertical is very close to one posting per pixel.
As the volume of postings on news.ycombinator increases due to increased traffic to the site the graph will stretch more further to the right.
This could be counteracted by changing the algorithm to 'bin' more posts to the right hand side to get for instance one month per bin, but in practice the outcome would be the same, you'd just have another weighting to do to get the Y-axis of the bins to line up.
Those are only available to logged in members directly from HN.
I agree that that would make it a lot better.
The biggest indicator of something being 'populist' but not 'HN' is when it gets killed after receiving more than 10 upvotes.
The tagging has been a lot of work, to put it mildly and it is far from finished. Eventually I hope to crowdsource that part to get it perfect.