Show HN: Updating archive of HN's /newest, with pre-rendering of the URLs
hnflood.com
hnflood.com
And I couldn't find any RSS feeds of /newest (no one that crazy I guess), so I wrote something to grab pointers to the submissions and render them as GIF, HTML, and text.
If this is useful to anyone I'm happy to improve it (including design).
Also, really dislike the target="_blank" on every link - I can get that anyway with ctrl-click.
Was surprised you're generating gif previews but not then inlining thumbnails?
Re: _blank - it was a peeve of mine that I'd accidentally miss a ctrl on the click reading /newest, but with this interface that's not a problem since moving back will work and keep your place. I changed it and the historic pages are regenerating now.
Re thumbnails, I wanted pages I could load quickly on ipad, gogo wireless, bad 3g, etc. Happy to do a version with thumbs this week if you'd like.
However some of the website owners may not like you're downloading the content of their articles and hosting them elsewhere. There is a great army of spam sites that will just watch http://pingomatic.com/ and scrape each new entry. Then they will on host it on a splog and stick adverts up. Which gets website owners annoyed.
Trouble is Google might not realize which is the main site and won't get the page rank or the visitors don't come to their site but an alternative. They'll get annoyed as they won't be able to monetize them or see who is reading the page through analytics.
So you may want to set up a DMCA page and abuse email address to stay on the right side of the law. Also a robots.txt which denies Google and Bing from crawling the pages you downloaded.
I'm not planning on copying any of the actual HN content, and don't present copy at all if it is on news.yc. At some point I'll hook into the API to grab comments/points every so often to update into the index pages and probably allow voting from the pages.
Very nice idea otherwise! I wonder if you could get in (legal) trouble for providing the screenshots/texts.
Great work!
Actually... It's just using the unix fs as a db right now, the structure is open and pretty easy to decipher. The dir tree for the objects is just "db/substr(md5hashofurl, 0, 2)/substr(md5hashofurl, 2, 3)/md5hash)" and the bytime/yyyy/m/d/hr/min is just symlinks to the md5hash (0 len file).
Usually I am just scanning to get the sense of whether it's worth reading every word, and if so I go to the original source.