Show HN: Hacker News Classics
jsomers.net
jsomers.net
It makes sense: on a site devoted to news, an article posted so long after it was published has to be especially good.
So I hacked together this page, which links to every HN post with a date in its title earning more than 40 votes. It’s sorted in chronological order to encourage wandering.
(dead) http://yorktownhistory.org/homepages/1900_predictions.htm ===> https://web.archive.org/web/20100108205037/http://yorktownhi... (archive from 6 days later)
I don't think this can be done programmatically though... thanks for putting this together. I enjoy the old posts a lot, too.
http://archive.org/wayback/available?url=http://yorktownhist...
It returns a json with, notably, the closest archived page given the timestamp.
http://yorktownhistory.org/wp-content/archives/homepages/190...
Ezra was a... complicated man, but he had keen insights.
I probabbly wouldn't understand most of them nearly as well if they weren't proven out by recent history ;)
Out of the 2k+ posts you list, some people are good at finding good quality classics!
The top 10 here combined accounts for more than 10% of all the posts:
USER POSTS COUNT
Tomte 66
luu 42
tosh 33
ColinWright 27
adamnemecek 26
vezzy-fnord 26
brudgers 25
pmoriarty 24
networked 23
shawndumas 22Well, let's try again: https://news.ycombinator.com/item?id=16444460
I wonder if you could ensure fewer false negatives (i.e. find even more great stuff) by doing the opposite: attempting to filter out every post whose link is to a page that came into existence within a month of the post's submission date.
This would likely require scraping the source links (unless you can get that from the https://cloud.google.com/bigquery/public-data/hacker-news dataset or somesuch), but it might be worth it anyway. It'd literally be "Hacker News, minus anything that looks like News."
And now I’m reading Arnold Bennett’s ‘How to Live on 24 Hours a Day’ originally published in 1908. Good to know tongue-in-cheek self-help books will never go out of style!
" The idea of devoting to them thirty or forty consecutive minutes of wonderful solitude (for nowhere can one more perfectly immerse one's self in one's self than in a compartment full of silent, withdrawn, smoking males) is to me repugnant "
Or click on the date in an account's profile to see what HN looked like on its birthday.
Out of curiosity, for the date rendering, did you xdef racket's date libraries or implement your own date algorithms in arc?
I guess you could've also shelled out to `date -r <timestamp>` which would get the feature done in about ten seconds.
Yes, and we specifically left it as an exercise for the pokey reader to figure out why.
Arc does its own date stuff. We extended it a bit to be able to print things like "x months ago". Edit: and "Feb 31, 2017".
We use it for a browser extension that helps a lot with moderation work. And for general experiments. If I do something more ambitiously new it will probably be in Arc.
Oho, so Arc does compile to JS now.
You gonna release the compiler or what? (Sometime this decade.)
Edit: should be fixed now.
You can't deny it anymore. Everyone now knows that 31st of February exists.
So downloaded the version he offers for download and am now installing the dictionary using the steps he provides.
Thank You, really looking forward to using it.
edit: and his install steps still work, using it in the dictionary app right now. Super!
I think that data could be very interesting, and would also serve as a sort of “hall of fame” of articles the HN community loves (moreso than just a list of the most upvoted articles of all time).
Of course, you could also figure out the optimal duration between successive posts, and figure out when to repost yourself, if anyone wants a karma-grab. ;)
The best example I can find on short notice is this hexagonal grids article:
https://hn.algolia.com/?query=Hexagonal%20Grids&sort=byDate&...
You could likely find more by searching HN comments for the phrase “previous discussion,” because people tend to post links to the previous HN submissions.
It doesn’t bother me, since there are likely always new HN readers who haven’t seen great “classic” articles.
The HN FAQ says:
> If a story has had significant attention in the last year or so, we kill reposts as duplicates. If not, a small number of reposts is ok.
Do you do some deduplication. Some classics are resubmitted a few times successfully. Perhaps add a third sort criteria, that is sorted by the number of big reposts.
For example say that several years ago a post might hit front page and 500 people would vote on it. If 300 people voted up and 200 down then it gets 100 points.
Fast forward and now there are more people on the site. That means greater confidence in the percentage but because only points and not percentage of up/down is revealed, a post with similar percentage up- vs downvotes appears to be better. 600 up and 400 down is still 60% up vs 40% down but it results in 200 points total.
Like I said though maybe HN does account for this and 100 points today means the same percentage of people upvoted it, I don’t know.
https://news.ycombinator.com/item?id=243417
As a result, its score did not increase, although upvoting "worked".
for (var i in data[yr]) {
var story = data[yr][i];
if (seen_titles[story.title]) continue;
seen_titles[story.title] = 1;
// ...add story to list of stories
};I created this for my own use but sharing it here because others may find it useful (hope this isn't breaking any HN etiquette and if you're the original author and want me to take it down please let me know).
[1] https://docs.google.com/spreadsheets/d/16MNPM9fhpglC1s1OV-tp...
[2] https://gist.github.com/san-kumar/0b7d218a08d4d8b4a7d9c07db5...
For security, ad-blocking, and privacy reasons I have javascript blocked in my browser, and really don't like to turn it on except for sites like my bank.
Could be done many ways.
Lets imagine you like text, you dont mind reading something somewhat structured and regular like json and you dont need all the html, css, etc. tags and window dressing.
5-minute quick and dirty solution
1. curl -4o 1.htm http://jsomers.net/hn/stories.json
tr , '\12' < 1.htm \
|exec sed '
1i\
<pre>
/./{
/{\"/,/}/!d;
}
/\"url\":[^{]/{s/\"url\":/&<\/pre><a href=/;s/$/>[FETCH]<\/a><pre>/;};
/\"objectID\":/{s//&<\/pre><a href=https:\/\/news.ycombinator.com\/item?id=/;s/$/>[FETCH COMMENTS]<\/a><pre>/;s/\"//3;s/\"//3;};
s/{\"created_at\":/\
\
&/g;
/}}}/s/$/\
\
-----------------/g;
$a\
<\/pre>
' > 2.htm
2.
Navigate to file:///2.htmWere there no pre 1900 refrences or is that just were you chose to start?
[1] https://en.wikipedia.org/wiki/Complaint_tablet_to_Ea-nasir
well done!
Perhaps Hacker Classics could automatically ensure all remaining live links get archived?
Until someone stops paying the host server or the domain expires and then those books just disappear.
Done. I'm going to run through every URL ever cited on HN at some point when time permits.
https://github.com/jsomers/hacker-classics/blob/master/index...
and I just reliazed one of the reason why classic writings are special -- especially for another writing material purpose is because it has passed the the test of time..
Check back a few hours later.
(Scraped content to markdown, simple Jekyll website in progress)
> if (data.hasOwnProperty(yr) && parseInt(yr) < 2012)
best
You could say they are… evergreen (title bar color)
[1] https://github.com/minimaxir/hacker-news-undocumented
[2] https://news.ycombinator.com/classic
Quite a cool coincidence!