Hacker Hacker News - see just the programming/math/science links from HN
hackerhackernews.com
hackerhackernews.com
I'd appreciate any feedback on this.
* News.YC, which rarely has articles about programming anymore (perhaps one or two a day).
* Programming Reddit, whose commentators think every article is a jab at The One True Idea, which they seem to think is PHP.
* Lambda the Ultimate, which only covers a small area of programming; language design and implementation
I want a site like LtU but with a broader appeal.
Implementation (!= Programming Language Theory) is off topic for LTU, unfortunately.
(There is overlap with the programming and web development, but something like usability guidelines fits in web development but not programming. By my definition, anyway.)
(I am sticking around because of the general interest articles which is what I used to visit reddit for before it became too popular)
I'd be curious to know how you're formatting the input data. If you're just using the title, all of the text parsed through BeautifulSoup, some parsing algorithm to obtain the body of the article, etc.
Not too long ago, I wrote a python script to automatically extract the bodies of articles, which you may find useful: http://blog.davidziegler.net/post/122176962/a-python-script-...
I've got to say, your classifier looks like it is amazingly accurate for those distinctions. I'd be really interested in seeing some of the very determinative words. (There are some obvious candidates like "Erlang" and "Twitter" but every time I do a natural language processing project I'm amazed by how our distribution of "little meaningless words" changes markedly between FeatureA and FeatureB.)
I've been wanting to use that library for something just like this. I would love to read a write up of this, or see the code at some point.
Thanks
edit: source at http://hackerhackernews.com/hhn-20090706.tgz
Can you specify the timezone?
Also, how often is the page updated?
Is it specifically subreddits to avoid? I completely agree that subreddits are a terrible implementation of user-content customization because they force items into a single-parent hierarchy.
However, categories that filter content according to different weights would be quite useful. By pre-selecting these categories and thus pre-computing them, it would even be reasonable to implement (actual per-user filters might not work so well in real-time).
The reason hacker news sucks less than other places is cause when people come up with cool mashups that could be interpreted negatively (you don't like all the news chosen by the community, whaaatttt?!), no one gets pissy or flamey.
It's almost like "the way society should work" or something...
The classifier is already struggling. It doesn't seem that this is a sustainable way of classifying links, especially since the classification of technical/non-technical is arbitrary itself. I'll be trying out some other things to improve it today; email me if you want to chat about it.
Are you training on comments as well? My bet is that the comments section will be more useful for this than the actual article.
Incidentally, I really wish that upvotes and the like were publicly visible. This would surely result in a similar tool, but for a recommendation system.
...so meta!