Websites That Feed Hacker News: Top Sources of Submissions by Median Score
github.com
github.com
NYTimes.com is certainly more prevalent but, just like Grauniad lately, it seems to be diluting quality on HN rather than being important. Same goes for Medium mentioned above.
Whereas sites like patio11's or cperciva's blogs, YC startups (bu.mp), tutorials are what makes HN unique and interesting.
I think this underscores the difficulty in quantifying the nature of "quality", especially for a broad audience. I generally check the NYT homepage every day, so seeing its URLs on HN isn't particularly helpful to me (ignoring the value of the HN discussions)...however, there is so much interesting information on a daily basis, period, that I bet if the HN front page consisted solely of the most-upvoted of high-traffic mainstream sites, e.g. github, nytimes, medium...it'd still be interesting to me because there'd be a lot that I would've missed otherwise.
That said, it'd be cool to have an option/Chrome plugin to filter the frontpage links to domains with relatively rare submissions, just to be able to quickly see the unique upvoted submissions for the day.
A little off-topic, but this reminds me of how non-straightforward it is to categorize what the true "domain" of a given URL. Blogspot.com has the most submissions by domain, but the API that runs the ?site query on HN returns just 5 links:
https://news.ycombinator.com/from?site=blogspot.com
I suspect that's because HN differentiates between somerandomdude.blogspot.com and googleblog.blogspot.com...but that in itself is an editorial/arbitrary decision...Why is subdomain more relevant for differentiation than for github, in which github.com/blog is grouped in with github.com/somedudesrepo?
And FWIW, it seems subdomain faceting is done manually...education.github.com links are shown on HN's page as just github.com...whereas googleblog.blogspot.com's domain is fully listed.
blogspot.com is in it. github.com is not, but github.io (where Github Pages are hosted) is. I would guess this is what HN uses.
That said, it of course can't help with e.g. categorizing github repo links by user, since those are by path rather than subdomain. Ah well.
After SeanDav's question and minimaxir's comment, I summed up reposts' scores before computing the mean and the median:
HN news sources by mean score: https://docs.google.com/spreadsheets/d/1tTDDG2xg7OVKdUy4WCZ_...
HN news sources by median score: https://docs.google.com/spreadsheets/d/1P20sKg-fI6msZVZtJFe0...
HN news sources by number of submissions: https://docs.google.com/spreadsheets/d/1mmfbNWaX0Nr1P65VmwZp...
SQL code: https://github.com/antontarasenko/smq/blob/master/hackernews...
Is the HN algorithm rigged in favour of things he writes, or does this community really get a lot out the things he says?
[1] http://blog.samaltman.com/asana
[2] http://blog.samaltman.com/the-tech-bust-of-2015 made me laugh, for example
"Startup kids" is too dismissive. Some of the very best comments about startups come from grizzled veterans. Will ChuckMcM or Animats mind if I call them grizzled? Let's just pause to appreciate what incredible value they and others add to this community from the wealth of their experience.
Than again, depending on your definition of "kid" there are "kids" on HN whose experience with startups is already impressive. Experience should perhaps be measured in iterations, not years.
HN has many subgroups, including plenty of hackers. Plenty of purely technical stories make the front page. And the startup and hacker groups overlap.
We get complaints about the balance whichever way HN trends.
Wrong. When I first lurked here a high-percentage of posts here were relating to startup concerns. The number have dropped dramatically over time. My non-hacker son-in-law who only knows finance was a hacker news reader a few years ago, hoping to learn tips about starting something up. He gave up due to the dwindling number of such posts.
FWIW, if anyone from that site/mag are frequent HN readers, HN is the reason I subscribed, and gifted subscriptions to several of my family for xmas this year.
There's a daily deluge of articles from ars, techcrunch, nytimes, etc, so the (tons) of articles that do get to the top get penalized by the ones that don't.
I don't think there's a flag for "hit front page" so might have to estimate that with a min point filter instead.
1. Sorting by the median. The mean is not very informative for the quality of the source. Most sources provide low-scored content with eventual hits that drive the mean up. The median fixes this problem.
2. Cutting off at 10 submissions. An arbitrary minimum to exclude pure luck from the results.
In the end, this ranking excludes websites like github.com and youtube.com, but it features some less known sources.
That said, I would be interested in the mean, just to see how different the two lists might be.
The 10 story minimum is to ensure a reasonable threshold for error and so a single submissions with 1000+ points (e.g. Show HNs) don't skew the results.
[0] https://docs.google.com/spreadsheets/d/1-TCo1mxiTkO4ZiXg5acU...
Even more, I'm starting to see more post from medium nowadays that has declining quality relative to 2015.
If Reddit and YouTube submissions can have ranking penalities due to highly variable quality, so should Medium.
https://web.archive.org/web/20100415161333/http://muckandbra...
In 2011, it was redirected to this blog, which is still live: https://cemerick.com/