Hacking Hacker News
joelgrus.com
joelgrus.com
Newspapers may be old-fashioned, but here's what we're losing if you never see one. They are trying to tell you what's actually important, not just what's important to you. You may not read the whole paper, but you at least see headlines, making you aware that something's going on outside of your microtargeted world of fashion or music or Wiccans or zombies or whatever you're into.
Replace 'newspapers' with hacker news and you get the point.
newspapers reported stories to sell you something, be it papers, ads or someone's agenda, not because they believed we'd all be more rounded citizens.
Whether that bubble is a subset of some other bubble, it's still a bubble.
I would love to hear why vanilla HN is better now.
1. The Hacker News API I was using was very unreliable and would go down for days / weeks at a time, which made the whole pipeline unreliable.
2. I was consuming this as an RSS feed, but when Google Reader shut down I abandoned my RSS habit cold turkey, so now I pretty much only read sites that I visit directly, or things people link to on FB / Twitter.
Probably because it has more politics and less 'hacker news' these days.
(Note: I usually prefer the "unfiltered" version, unless there is a very special new that covers the 90% of the front page.)
However, I haven't completely given up on the idea. The point is that while limiting news to things you've liked in the past is a terrible idea, helping you find good articles isn't.
The question is, what constitutes a good news item? I think the key is that there are many answers to this question:
1. Something very specific to your interests (your classifier should be good at detecting this)
2. Something big that everyone should hear about (the HN frontpage is good for this)
3. Something about a new concept you've never heard about. I've implemented this by keeping track of popular words on reddit and giving a boost to words not seen in the past, with some success.
4. Probably others I haven't thought about.
I'd love to hear your thoughts on this.
If you let your program log into HN using your account, it should be able to tell which of the stories you've up-voted there. If you use that as the input to your classifier, as you read stories on HN, simply mark those that interest you by up-voting them.
I'm also curious to know whether the stories are weighted by age to account for changes in what you find interesting.
[1]: https://chrome.google.com/webstore/detail/hacker-news-sideba...
Also, you can set hitsPerPage = 1000. ;)
"you can only fetch the 1000 hits for this query, contact us to increase the limit"
It was a nice try, but it did seem too good to be true.
Just pass created_at_i<X, where X is the time stamp of the earliest submission.
I was able to download 500k stories (i.e. about half of HN's 1.26M stories) before I ran into memory issues; I've fixed them and am downloading the rest.
- Python module: https://github.com/karan/HackerNewsAPI
- REST API: https://github.com/karan/HNify
RATE LIMITS
We are limiting the number of API requests from a single IP to 1000 per hour.
If you or your application has been blacklisted and you think there has been
an error, please contact us.You can query 1,000 stories per request, and do 1,000 requests per hour. That's 1M requests per hour. There are 1.26M Hacker News stories indexed by the API. :)
EDIT: Finished downloading all the entries (and can confirm that 1.26M is indeed all of them). Took 3 hours due to a conservative wait period between each request to make sure I stayed within the limits.
I personally despise recommendation / personalization algorithms of any kind. I still have never found one that's actually better than myself at distinguishing articles that I'd like to read, music that I want to listen to, tweets I'd like to see, etc.
When reading HN, I'm constantly surprised by links that would not normally be on my radar for things I'm interested in. I think personalization algos, in general, are good at filtering those away.
Since the author mentioned HN being too much of a firehose and this then also being a solution to the "too many links to keep up to date on" problem, the solution might be a bit simpler than the author suggested.
HN already has the "best" links at https://news.ycombinator.com/best
It's hard to find - it's in the 'Lists' section in the footer, but it's still there and I use it all the time, when I haven't been actively reading HN for a while.
- WashingtonPost
- BusinessWeek
- MarginalRevolution
- NY Times
[1] Based on BuzzSumo's social data:
http://app.buzzsumo.com/#/influencers?q=@joelgrus&type=influ... (Press View Links Shared, Analyze Links Tab]
DataTau (the HN for data mining) seems to have failed, so I imagine a filter is the way to go rather than make a new website.
This raises the question: How does YC justify hosting costs? My completely-off-the-cuff-assumption-take-this-with-a-huge-grain-of-salt is that YC benefits by having a huge audience to make announcements to, like job postings at YC funded companies, various pg essays, or just investing in overall goodwill from the HN audience. Probably the most likely reason is to increase deal-flow to YCombinator itself, though.