Hacking Hacker News
joelgrus.com
joelgrus.com
If I (jcr) want to know what classifies as "interesting to Joel" (joelthelion), I simply look at the "comments" and "submissions" links in your HN profile. It will show me the stuff you took the time to comment on, or took the time to submit to HN.
If I want to know what's "interesting" to me, the "saved stories" link is visible in my own profile even through it is not visible to others. In the "saved stories" is THE goldmine of every submission I've either submitted or up-voted.
https://news.ycombinator.com/saved?id=jcr
Depending on your personal bookmarking habits, your bookmarks file/db can be another useful input. I'm in the habit of bookmarking both the submitted article, and the HN discussion page (if it's good). Assuming I didn't book the HN discussion, I could easily find the HN discussion/submission for all of the sites that I've ever bookmarked with a search engine. Give google an URL and ask it for all of the sites linking to that URL, then parse for HN, and you've got the target.
The serious problem I see with your approach was already mentioned by ck2; you're creating a bubble and will miss out on all the fantastic stuff that is interesting to you, but you don't yet know that it's interesting to you.
One of the primary benefits of HN and similar sites is learning about the things that others find interesting. Those things may not interest me, but the fact that others find them interesting is, well, interesting.
Why to they consider it interesting?
Why do I consider it uninteresting?
Even if my personal opinions remain unchanged, these are important questions for me to keep asking myself, repeatedly.
Personally, I think I'd be perfectly happy with an old-school killfile: do not show me posts whose headlines contain the strings "X", "Y" or "Z", or that link to sites "A", "B" or "C".
Or maybe even going down the "subreddit" route, i.e. a separate section for Apple content, Python hacking, SOPA/PIPA etc. In other words, create a "bucket" for the general high-level, recurring HN topics and allow users to pick which ones they want to see and those they don't.
Make two classifiers, one that tries to decide "is it interesting for thristian". And make another that decides "is this similar to an existing story". Then it gets demoted only if both it is not interesting and it is similar to something that already exists.
The problem with that is these days there's so many "clickbait" headlines, some of which offer absolutely no insight into the content itself.
Without policing each and every submission to ensure the headlines accurately define the linked article - it would be difficult to guarantee any kind of success with basic headline filters.
It was scrapped and eventually subreddits were introduced. I think people like the idea of communities.
The question of whether this becomes an echo chamber and you stop finding things that are interesting but outside of your current tastes is deep. There’s an old saying, “The best present is something you didn’t know you wanted until you unwrapped it."
The question is, how do you figure out what's interesting to who? Joel is attempting to answer this for himself. A site that could answer it for everyone while still somehow maintaining social bonds would do very well I think.
I'm still working on recommendations. There are suggestions once you follow someone. It's a good way to get started, but the best way to get good content is to hand pick folk based on what they have shared.
My theory is that as communities grow, you have one of two choices to maintain cohesion: aggressively police who is let in or give people with divergent interests space to express those interests.
If you don't, and your community is growing, eventually the most valuable members will no longer find value in the community. This is just reversion to the mean. But the consequence is that those valuable people will then leave, and the cycle repeats until the community evaporates entirely.
Subreddits are a "pressure valve" that allow high-value users to self-select and maintain their own set of norms without sacrificing the growth of the main public square.
It's not the stories and comments that are fundamentally valuable to HN, it's the people.
I remember about 5 years ago when I was looking for something else like reddit only more technical, saw HN and thought, "wow, I don't have a clue what 90% of these people are talking about. Seems interesting."
HN's subtitle is "Links for the intellectually curious"
I guess HN has a dual purpose to keep people up on "breaking" hacker news, but I like to think of it as "hacker news outside your thinking pattern".
Also why do people immediately go to AWS for testing something? Doesn't a real hacker have their own server handy for experimental projects, or is it only me?
That's something I think could be automated (at least for me.)
I can (and do) guess whether I am likely to be interested in a new programming language, based on what other languages the commenters compare it to.
In my case it would probably be something like:
* Compared to Haskell, Io or SETL => very interesting
* Compared to Java, JS or Scheme => neutral
* Compared to Groovy, PHP and VB => uninteresting
What makes you think that a Naive Bayes classifier will automatically put all unseen/unexpected/surprising data in just one of two categories?
Indeed, to the classifier it's just two categories, it doesn't "know" what's "in" or "out".
In fact it is much more likely that the classifier will sort of distribute the data that is unexpected (read: doesn't contain many features that are trained on) rather evenly among the two categories based on other features. Which is exactly what you would want it to do.
The author also said he plotted the precision-recall curves? (too bad he just showed a screenshot with the numbers instead of the graphs) That sort of analysis is bound to bring out such behaviour.
As other people have pointed out, the naive Bayes model works topically, so it will learn that I like stories about "patents" but not (usually) whether the stories are pro or anti-patent. It is totally true that I might miss an interesting story about the new OSX or about Pinterest, but I'm willing to live with that.
Two larger points are that
1. HN is only a small fraction of the news I consume, so it wouldn't matter that much to me even if it were a bubble chamber, and 2. The main reason I did this is that I simply couldn't keep up with the volume of stories otherwise.
Last night when I spoke about this, someone asked me whether I was concerned about all the false negatives I was missing. But before I started this, my RSS feed had like 800 (and growing) unread HN articles in it. Reading some of them, even a targeted some, is better than none.
Anyway, thanks for all the comments. I'm suprised (and glad) that people are so interested in this!
One of the nicest aspects of it is that it doesn't support a user's confirmation bias: your perspective isn't taken into account in filtering since it's just looking at keywords. That's probably not as important here, but especially on political news it's highly relevant. If I'm a Democrat, I don't want just left-leaning news to come through the grapevine, because that prevents me from seeing the other perspective.
The latter is a lot harder to extract from a bag-of-words. In fact the only thing I can imagine is if an article uses a lot of euphemisms or negative synonyms of a topic.
So you get the factual reporting articles regardless of left/right bias. With the more polemic hyperbole using articles, it will filter a bit more of the angle you disagree with, which adds some bias, but it also good for your blood pressure and a polemic article you disagree with is not going to make your view more balanced either.
Edit: Found it - https://github.com/joelgrus/hackernews
The model can only get better with more training data,
which requires me to judge whether I like stories or not.
I do this occasionally [using] the above command-line tool,
but maybe I’ll come up with something better in the future.
Well, you could analyse your server logs to see which stories you really did click on and which you skipped over (You'll probably want to only consider pages that have at least one click for the edge case of "I didn't even look at that page". Need a cookie or login to make sure it only counts your clicks)Also consider scraping your HN "saved stories" list as a positive source.
Don't recall if you mentioned it in your article, but you'll probably want to randomly insert the occasional low-scoring article as a check for under-weighting.
I really wish our profile pages supplied a (private) log of up/down votes for comments and flags for stories in addition to the the up votes for stories. It would make for some interesting datamining
Beyond that, only pg could say...
I'm working on something right now and this limitation causes difficulties. (http://cl.ly/241x3F3d123P0w0L1s3q)
because code written by N hackers scraping HN has a greater effect on the site, compared to using the reliable HNSearch API ~ http://www.hnsearch.com/api
Ask HN: Does HN have an API and if not what's the etiquette for scraping?
http://news.ycombinator.com/item?id=2138730
Ask HN: Is there an API for HN?
http://news.ycombinator.com/item?id=1107874--
I had a go at scraping HN myself a couple of days ago (in PHP) and here's the output:
http://thequeue.org/api/frontpage.xml
http://thequeue.org/api/new.xml
http://thequeue.org/api/best.xml
The trouble is that sponsored posts, i.e. We're Hiring (YC 11) break the code. If anyone can help, there's a question on Stackoverflow about it:
http://stackoverflow.com/questions/9301215/scraping-hn-front...
Hopefully when I get back to hackernews sometime tomorrow this will be frontpage where it belongs. :)
Nevertheless, I believe that a good news source, like a good community, sometimes gives you things that you might not like--things that challenge the filters we already have.
It's what flipping though a regular newspaper does, what hacker news does, what listening to a good broadcast does. The key is if you know the content is going to be so good that you're willing to take that risk despite your hesitation.
Hacker News is a bit like a firehose, but that's why I read it. The people on this site are informed, opinionated and pretty damn smart. It reminds every time I get on how little I actually know.
To be honest, a small part of me feel uncomfortable with that but that's why I read it--it's news I need to know.
I'm considering making the project open source as well so technical folks can run their own versions and modify it. Sort of a wordpress model where it's open source but they still make money off the hosted version which is simpler for non-technical people to use. Thoughts?
Even though it's not always fair, even though you have to wade through stories, the fact that there is a common home page is what spurs the discussion and keeps things interesting.
This community is exactly how I remember digg in its heyday. Mostly tech stories and one common home page that has some bit of prestige when your story reached it. That's the magic sauce.
I have this theory that "atemporal" stories (technical analysis, insightful essays, etc) that keep getting resubmitted every year are more interesting than news about the latest gadget. I've written about it before: http://news.ycombinator.com/item?id=2505081
I think there's probably some really interesting data in the comments (maybe just take top ten comments, 3 layers deep?) And wrt following links, one idea that occurred to me was taking a screenshot of the page linked to and basing part of your model on that... I suspect there may be graphical, layout similarities in some of the pages people like.
But kudos for doing it, very cool!
Looks like a great start to an awesome project. I think the next logical step would be to expand to give anybody a filtered HN experience, but you probably didn't need me to tell you that :)
I came from Java background where Maven reigns supreme when it comes to build + dependency + convention on file structure and I like this set-up.
What is the equivalent to that in Ruby, I know there's Bundler and Rake, but Rake feels like Ant where you'd have to do a few things yourself.
thus the more novel the content - the higher up it will remain
unless it's an article about a novel completely revolutionary arrangement of old ideas which you never liked on their own :)
My biggest complaint with the number of stories on HN is not finding the great articles but wading through a bunch of stories about the same three or four topics I am not interested in...
Apart from the technical side, I find the light blue text hard to read.
Just like you can do on Reddit.
I filter HN with http://hacker-newspaper.gilesb.com/, which pulls RSS, filters it, and reformats it on an hourly cron job. I mainly did it for the typography -- I disagree with just about every visual design decision on Hacker News -- but added very primitive filtering after the fact. I throw out any story from TechCrunch, Zed Shaw, Steve Yegge, and Jeff Atwood, because I just got tired of them, and any story with "YC" in it, too, because I got tired of seeing job ads for Y Combinator startups. (In fact it was the job ad for a Curebits marketing manager right after their scandal that did it.)
When Apple launched the iPad, I went in and added a simple regex to filter out any story about it. Hacker News is a great source for skimming but occasionally gets fixated on topics. I get like a hundred uniques a day so it's not exactly a huge hit, but I've thought about making a commercial version with customization. It got featured on Mashable and somebody created an iPad app which looked very, VERY similar, which I'm going to take as validating my design. But whether or not I ever startupify it, anyone who wants a customized version can just fork the project, deploy their own version in like ten or twenty minutes, and tweak regexes to their heart's content. It's on GitHub (https://github.com/gilesbowkett/hacker_newspaper) and only requires the most basic proficiency with cron, ruby, and python.
I also want to add comment-scanning. Right now I don't use comment links at all. The code extracts them but then simply throws them away. I don't want to add comment links back in unless I can also set it up to alert me if the comments thread contains comments from raganwald, patio11, jashkenas, amyhoy, etc -- basically automated comment elitism. I'm not trying to be a dick with that, I'm just a busy dude.
Anyway, when I set out to do this, I planned on doing a bunch of Bayesian whatnot, but I found that I got most of the way there just tweaking regexes occasionally. Likewise there's a lot of rough edges I could clean up, e.g., text encoding is a bit of a mess, and summarization in the style of http://tldr.it/ would make it way more useful.
But I recommend it because making deliberate decisions about what info you want to get from HN makes it a lot less like watching TV and a lot more like doing actual research into topics which interest you. It's surprising how much more enjoyable HN becomes when viewed through a customized filter.
https://github.com/gilesbowkett/dotjsfiles/blob/master/news....