Hacker News API
blog.ycombinator.com
blog.ycombinator.com
I count 8,483 submissions. I'm sure there's something interesting to be done with all of this data. A word frequency chart?
---
Edit: So apparently there's a ruby gem that lets you feed it a body of text and generates pseudo-random phrases based on that text.
I present to you the patio11 impersonator: https://gist.github.com/christiangenco/e8d085e47479be0131e1
One of my favorites:
A nice set of challenges -- kitty at a school with tens of thousands of bucks a year or less immediately.
---Also, a word count on patio11's submissions: 1,052,351. For comparison, all 7 Harry Potter books total 1,084,170 words. patio11 has written the entire Harry Potter series worth of content on HN. Just... wow.
Yesterday, I published an analysis on all Hacker News comments: http://minimaxir.com/2014/10/hn-comments-about-comments/
There's a lot of interesting trends in the data. Let me know if you want to know anything in particular and I'll get back to you. :)
I also included the number of thread wins (user made comment with highest number of points in a given submission thread) to see if there was any unusual relationship. There wasn't: correlation is 0.89. I ended up not including it in the article because it's long enough as-is. :P
SQL query:
SELECT
author,
SUM(num_points) AS total_comment_karma,
MAX(created_at) AS last_comment
FROM hn_comments
GROUP BY author
ORDER BY total_comment_karma DESC
LIMIT 1000https://hacker-news.firebaseio.com/v0/user/tptacek.json?prin...
I wasn't going to comment, but "mhartl" looked familiar. I just wanted to say "thank you" for your Rails Tutorial; I don't think I would have ever learned to program without it. It literally changed my life.
Click on the line chart to do an hnsearch for the time period.
Update: Site should be back up. It crashes occasionally (that's part of why I hadn't post it yet.)
The book is copyrighted by Ed Weissman, so although it's possible "Ed" is short for "Edna," the higher probability can be assigned to "his" in this case.
For folks who want to do interesting things with the API but don't want to be abusive to Firebase's servers, I whipped up a quick ruby script to cache a particular user's comments/submissions on disk: https://gist.github.com/patio11/1550cad3a02edd175049
It tries to rate limit itself by putting 200ms of sleep between requests, so downloading all of my comments would take ~30 minutes.
"I release this work unto the public domain." -- feel free to adapt it to your needs.
Usage is "ruby slurper.rb $USERNAME $MAX_COMMENTS_TO_FETCH."
From https://www.firebase.com/pricing.html it looks like the top plan supports 10k concurrent connections, so I suspect the impact is negligable.
Thanks for being an outrageously good resource and beacon of inspiration! You've unknowingly been one of the most influential role models in my career/life: I just relaunched one of my side projects as SaaS last month and it's succeeded beyond my wildest expectations (already at ~$8k YRR). Hopefully I can follow your trajectory and never have to actually work another day in my life :)
Though I don't know if I'd describe my lifestyle as "never having to actually work another day in my life." It feels less like work some days and more like work others. For example, it is 1:30 AM and while I could be snug in my bed I am instead clearing out the AR support inbox. (Poor planning earlier today, but still.)
It's certainly difficult at times, but when you can set your own hours, do wherever you want on whatever you want, and take as much time off as you want for any reason, it's difficult to justify using the W word.
https://play.google.com/store/apps/details?id=com.airlocksof...
The newest version (still under development, probably a month or two from release) adds support for displaying polls, linking to subthreads, and full write support (voting, commenting, submitting, etc). I'm fine with switching to a new API (Square's Retrofit will make it super easy to switch), but without submitting, commenting, and upvote support I have to disable a bunch of features I worked really hard on. Also it would've been cool to know this was coming about 3 months ago so I didn't waste my time.
Anyways, quick question on how it works -- when I query for the list of top stories
https://hacker-news.firebaseio.com/v0/topstories.json?print=...
it just returns a list of ids. Do I have to make a separate request for each story
https://hacker-news.firebaseio.com/v0/item/8863.json?print=p...)
to assemble them into a list for the front page, or am I missing something?
So, @anyone involved with the API project, can you give us an estimate on when will the OAuth-based user-specific API be rolled out? I'm fining with pausing my efforts until then, if it's going to be soon, in order to go a less complex and error-prone path.
[1]: https://github.com/geomaster/hnop/blob/master/backend/src/hn...
If you're on a supported platform, the Firebase SDKs handle all this efficiently and can even provide real-time change notifications.
If you use the SDKs, we handle the connection and all of the data is sent over a reused full duplex socket rather than individual requests. https://www.firebase.com/docs/android/, https://www.firebase.com/docs/ios/, https://www.firebase.com/docs/web/
I've never used the Firebase API itself before. It's very clean!
Edit: I reached the same (now obvious) conclusion as mentioned in the reply below. Now my quick hack is working perfectly. Thank you so much for this!
Cheers
Re write access and logged-in access, if that turns out to be how people want to use the API, that's the direction we'll go. But we think it's important to launch an initial release and develop it based on feedback. There are many other use cases for this data besides building a full-featured client: analyzing history, providing notifications, and so on. It will be fascinating to see what people build!
It would help me out a lot if the current front end would live on under oldnews.ycombinator.com like that until the new API has write access, though. I think it's pretty cool to be able to be reading an article somewhere else, click "Share" in Android and have "Submit to HN" pop up in the results.
This definitely does suck. I feel your pain. But it's also part of the package of scraping websites. You go in knowing that it could break at any time.
Thanks very much guys!
HN is definitely an example of a site that isn't ideal in a mobile browser. For instance, if you have the ability to downvote, it's incredibly easy to mistakenly downvote when you mean to upvote because how close and small the buttons are. There's other added functionality, like tracking who I've upvoted / downvoted in the past as well as tracking un/read comments when returning to a thread. In the browser, I use a chrome extension for this, and on my phone, I use airlocksoftware's app. (side-note, I wish said state carried between the extension and the app)
The developers of HN are surely capable of creating a mobile website that could work just as well, or even better than a mobile app. But currently, it's not ideal. And for that reason, I completely appreciate airlocksoftware's (and the devs of other HN apps) for their efforts.
Yep, there are no API methods specifically for getting the Show, Ask, New, Comments etc. lists, yet.
They will probably add such, though.
On the one hand, of course it's great that HN is finally getting a proper API and also modernizing its markup (which is a mess even if you ignore all the tables – for example, the first paragraph in a comment usually isn't wrapped in <p> tags), but on the other hand this current v0 version is very lacking and impractical for a regular client application.
Since the top stories (limited to 100) and child comments are only available as a list of IDs a client app would have to make a separate HTTP request for every single item, which is obviously not something you'd want to do especially in a mobile environment. Other lists apart from the top stories (new, show, ask, best, active etc.) don't seem to be available at all right now.
Of course this is just the first version, and the documentation promises improvements over time – which I don't doubt at all – but there's no clear indication that the API will be at feature-parity with the current website, even excluding anything that requires authentication, by October 28. So this means that I – and other developers of client apps or unofficial APIs – will probably have to write new scraping code once the new rendering engine (which I assume refers to the website) arrives instead of being able to switch to the new API immediately.
Now I guess I might just be needlessly worried, especially since the blog post explicitly says that the new API "should hopefully making switching your apps fairly painless", but then why not wait until it's actually ready for that before making the announcement? Putting a half-baked API out there a few days/weeks (?) in advance before it's fully fleshed out doesn't seem all that helpful, at least to me.
Oh, that's in the actual API.
https://hacker-news.firebaseio.com/v0/item/8422922.json?prin...
I strongly hope they add a plain-text field for returned comments.
I think it's even easier now that JS is a supported app language.
With access to a Firebase SDK the only major additions the API needs to become a viable replacement for existing read-only client apps would be support for all the other lists apart from top stories (new, show, show new, ask, jobs, best, active) and more than 100 items for each. For apps that need write access I'd suggest keeping the current website on a separate subdomain until that is implemented into the API.
EDIT: After having played with the JS SDK a bit I'd like to add that it is indeed incredibly awesome.
http://www.windowsphone.com/en-us/store/app/hacker-news/a527...
Here's a list of issues that I encounter very often:
- It often crashes during startup
- There are massive encoding issues, I see question mark symbols pretty often.
- Comments often won't load, without any error message whatsoever.
I uninstalled the app multiple times for these reasons, but unfortunately there isn't anything that's better.
queryparams: pageSize (no of items to return in one go) before (takes an epoch timestamp and returns results before that timestamp)
this is a stream of HN stories that have made it to the front page, in order of the time that they made it to the frontapge.
pip install hackernews-python
Usage: >>> from hackernews import HackerNews
>>> hn = HackerNews()
>>> hn.top_stories()
[8422599, 8422087, 8422928, 8422581, 8423825...
>>> hn.user('pg')
{'delay': 2, 'id': 'pg', 'submitted': [7494555, 7494520, 749411...
>>> hn.item(7494555)['title'])
Hacker News API
>>> hn.max_item()
8424314
>>> hn.updates()
{'items': [8423690, 8424315, 8424299...], 'profiles': ['exampleuser',...]}
https://github.com/abrinsmead/hackernews-python Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "hackernews.py", line 23, in top_stories
return r.json()
TypeError: 'list' object is not callable
Tried in both python 2.7.3 and python 3.2.3EDIT: You need a relatively new version of requests for this to work. The version packaged with Debian Wheezy is too old. Use pip.
Prior to version 1.0.0 response.json() was a property, and was not callable. That's probably what caused your error.
I'll add a minimum version to requirements.txt for the folks that install this package manually.
But more importantly, why no favicon? ;)
Good work.
(As I find myself pondering the idea of standing something up like this on an dual-stacked server purely so that I could access HN from my IPv6-only test network... hmmm...)
EDIT: After looking at the documentation there are two new aspects of the Firebase API not in the Algolia API:
1) Ability to see deleted/dead stories.
2) Endpoint for user data.
Question to kogir/dang: Has the "delay" field (Delay in minutes between a comment's creation and its visibility to other users) always been there?
From the hnsearch about page:
"Every 2 minutes, the last 1000 HN items (stories, comments, polls) are sent to Algolia's indexing API. Items from the last 48 hours are refreshed every hour."
Regarding the REST APIs, let's keep both for now :)
Example: https://github.com/minimaxir/hacker-news-download-all-storie...
Repeat.
Is HN data already in Firebase (as its primary data store) or is content from HN's DB getting 'mirrored/cloned' on-demand to Firebase for the API?
Array.prototype.forEach.call(document.querySelectorAll('a[href^=user]'),
function(v,k) {
var s = document.createElement("script");
s.src = '//hacker-news.firebaseio.com/v0/user/' + v.innerHTML + '.json?callback=ud_' + k;
document.head.appendChild(s);
window['ud_' + k] = function(user_data){
var avg_karma = user_data.karma / user_data.submitted.length;
v.innerHTML += ' (' + avg_karma.toFixed(1) + ')';
}
}
);It lets you download all comments/stories for a user as a JSON or CSV file, breaks down karma between comments and stories, and plots comment/story counts, karma, etc. over time on a line chart (clicking will show you the details via an hnsearch).
Also I built some npm modules so you can get this information via the commandline.
Example: http://hnuser.herokuapp.com/user/tptacek/
The Chrome extension hasn't been updated for awhile (it just superimposes a small amount of this information on the user page).
I know you can't not iterate because people are scraping, but it does stink. At least this will make everything more future-proof going forward.
However, it may be nice to give a bit more heads up than 3 weeks. I know a lot of apps can take ~2 weeks to get through the review process for iOS.
If people need help converting their apps, I believe the Firebase guys have offered to help. Contact us at api@ycombinator.com and we'll connect you.
I've built a library for iOS (https://github.com/bennyguitar/libHN) that handles scraping, commenting, submitting, voting, etc pretty well and allows me to make as few web calls as necessary to use HN. It looks like I'd have to drop functionality and completely change the networking scheme to match this API - something I'm not willing to do yet.
Correct me if I'm wrong here, but to get every comment on a post, I'd have to recursively get each item for each child. Instead, right now, I can make one network request and get all comments for a story. Granted, I have to parse the HTML (which I hate), but it's a much cleaner solution than going through every item, checking the children and then getting those items ad infinitum. Again, I just glanced over the documentation, but that seems untenable to me.
"Most importantly, the reason we released an API is so that we can start modernizing the markup on Hacker News. Because there are a lot of apps and projects out there that rely on scraping the site to access the data inside it, we decided it would be best to release a proper API and give everyone time to convert their code before we launch any new HTML."
- More than 100 stories
- Best, Top, Ask, ShowHN, Jobs, User Submission Posts
- User Management (logging in/out)
- Commenting
- Submitting
- Voting
My app doesn't just function as a reader, which is what this API seems geared towards with the v0 release, it functions as about as full-fledged of an HN client as you could get. There's a couple things that I haven't built in yet like changing your about me text, but those were on the roadmap.I'm actually thinking about storing the configuration of how my app scrapes online such that if the HTML markup changes, I won't have to push huge sweeping changes to the App Store to get my app online again. I just deploy to Heroku and the app will handle that configuration and scrape correctly sans pushing to Apple.
And that returns only the ids, nothing else. To get basic information like the score, title or url you have to lookup the ids individually. And even the story items do not contain such basic information as the number of comments. And you can't calculate it yourself since only the top comments are even returned (as ids of course). So you'll have to recursively dig through the comments to get the number.
This is even more curious as there is a very solid Algolia API where you can filter for submission time, story score, number of comments and even return a greater number of results + access page numbers to get even more.
To get the information of a single algolia api call you will need hundreds or thousands (in case of nested comments) "official" API calls. Hoping for updates
Right now one team, Ycombinator, is trying to fix important issues in the ranking and moderation of posts and comments. Many of us are frustrated by the increasing domination of popularity (and hatred) over quality and relevance. A lot of good submissions and comments are simply buried, never to be found. There is too much muck to have to wade through. The timing of posts and comments plays a much larger role than quality. I could go on and on.
Imagine a Netflix Prize-like flowering of experiments and collaboration, leveraging the hacker community's collective smarts and enthusiasm. Many of us have ideas, but right now are unable to test them. What a shame if a great idea dies on a notepad.
There are two possible issues with opening up voting data: gaming and privacy. If having vote data allows someone to game the front page, then only include it with some delay (2 days?) so that it could't be used to game the front page. This will still allow experimentation with collaborative filtering algorithms and the like.
My take on the privacy issue is that anonymity isn’t that important for a site like Hacker News:
1. Startup culture is about straight talk, putting your money where your mouth is, and open critical feedback, both in the giving and receiving. There are precedents for exposing voting data (e.g. Quora, Facebook, Stack Exchange).
2. HN is not aimed at political discussions or other topics where anonymity can be paramount.
3. Pseudonymity is sufficient for those who don’t want their votes and comments tied back to their actual identity.
Thoughts?
I would love to hear from others who yearn to experiment with alternate algorithms and strategies for improving Hacker News.
Even though it's read only, I'll continue to use my scraper rather than the API simple because it's one request, rather than the API would require one request for the top IDs and then one call per story, so it would be 31 calls instead of just 1.
Unless I'm missing something, it seems fairly poorly designed for top stories, and non existent for new stories.
------
EDIT: Looks like I missed the text about updating to a new rendering system in 3 weeks time, and to iterate designs faster to allow mobile friendly theming. Looks Like I WILL be updating to use the API
There's a bunch of css pages that come out for hacker news, but I couldn't find anything that aggregates them. This will be alot easier to extend and customize the site.
I'm not seeing any api's for the jobs or show sections though? Hopefully this might come in the future?
http://insin.github.io/react-hn
I've just gone for it with Firebase's React mixin, binding everything as an object, since their devs in this thread don't seem concerned about rate limiting. The mixin seems to throw an error every time I try to unbind, which I'm just catching and logging for now.
Edit: I just watched this comment pop up live in my version - pretty neat :)
Is the source available somewhere?
https://github.com/insin/insin.github.com/blob/master/react-...
Is there any chance of getting more than just the top 100 stories returned? I think it will be a lot more useful for api consumers if you can use a query parameter to set the limit (within reason, usually 1,000) and a number of results to skip. For now, scraping is still more desirable to me since I can retrieve any number of results in their current order.
Better yet, but more complex: a number to skip and a certain timestamp so I don't see the same article on two pages due to natural upvoting, downvoting, or rank decay.
Also, if there's any flexibility still with property names, I'd suggest these changes for clearer semantics: "deleted" -> "hidden" (since they're obviously not deleted) "by" -> "author" (for more clarity) "kids" -> "children" (the common convention)
For example, a site where HN members can upvote and rate different development tools, libraries, IDEs, management tools, etc. All with backlinks to HN discussions. It's a great community and there are many ways we could share knowledge and experience.
I'd rather use the REST API directly, for what I need is rather simple and not downloading, installing and maintaining an SDK is more appealing. (My app was developed a while ago and was doing HTML scraping, but the 30-second limit on HN killed it, because of testing -- I don't need to query more often than twice per minute, but while testing I ran the thing a little too often).
So, what are the limits on the REST API, and how do limits work? (A max number of requests per hour would be better than per minute for example).
1. Provide a way to bulk download the data (that's a click, instead of scraping the API)
2. Add a field for the maximum position a story reached on the front page
3. Add the numerical score for the comment (at least on comments that are N days old, which won't interfere with the reason to hide the scores on the main comments)
Some other changes that would be awesome (but are less realistic) include:
4. A historical event log of votes (even better would be relating those votes back to users, but I imagine that's not going to happen for privacy reasons. An intermediate possibility would be a vote log connected back to anonymized user ids, assuming the anonymous id -> real user id mapping is difficult)
5. A historical event log of display position changes for stories & comments
6. An event log of pageviews with as much metadata as possible to release without infringing on privacy
This is bigger news. No more tables!
1. Major functions are dropped (best, new, job, ask,...)
2. Useless response schema of the top stories API (should learn from Netflix's internal/public API design)
3. A short transition period
It's very hard to provide same responses to apps.
It will be interesting to see if it has an impact on site traffic, how much of that traffic is scrapers today?
I see the topstories query to replace scraping /news: https://hacker-news.firebaseio.com/v0/topstories.json?print=...
And there must be a query to get the newest items instead of the current /newest.
Are there also new equivalents for /active, /best, /classic, /show, and /shownew?
I'll be happy to replace dodgy XPATH parsing with a proper API. Hope we won't lose these other views though!
I'll need to "convert" SwiftHN (https://github.com/Dimillian/SwiftHN) either to this new API or adapt my scrapping engine to the new site layout.
The nice thing is that HackerSwifter public API is already in a quite finished state for the available functions, even if I switch to the API, the method calls will be the same.
I'm going to start diving into the API to build a simple, powerful "Google Alerts for HN" app on Assembly, and I'd love help from anyone who's interested: https://assembly.com/hn-monitor
There are some products like this out there, but they had to rely on scrapers and the HNSearch API, so I've always found them to be spotty. I think we can make something better.
My only problem with the API is that you need to do an awful lot of Ajax calls just to get something out of it. The topstories endpoint just gives you an array of IDs and then you need to do one Ajax call for each ID to get the story.
Oh well, the site is a work in progress. Not done yet.
https://github.com/stevenspasbo/Hackernews
EDIT: I should say, I started on this to create word clouds, but if anyone has any ideas, contribute!
Anyone with a Firefox phone that wants to work together on a HN client for it? FFOS has been lacking a good HN app.
I wonder how many other people have this? And what their techniques are?
On the top of my wishlist: Look up HN story by URL.
(From the comments, I learned about a different API that will suffice for now: https://hn.algolia.com/api/v1/search?query=google.com&restri... )
Now I will have to remove that. Please keep current version of the YC site on a different subdomain.
By the way, animations (the slide between list and comments) have been jerky since a recent update. I'm using the Android app on an HTC One M7 with the latest stock ROM (4.4.3 and Sense 6.0). A few other people have mentioned this in Play Store reviews. Is this a known issue?
Thanks HN!
EDIT: I just read the post about using the Firebase SDK to do this efficiently.
oh.... wait...
As I'd like to do some analysis, but don't really want to thrash the API, downloading all data.
One question: why the choice of returning everything as an ID, on mobile, this will require a lot of very small network requests.
They should really update to a modern web framework at the same time. Big modern frameworks like Rails are making ridiculously awesome improvements like replacing page loads with XHR (quicker loading since JS/styles/etc is all loaded already, no screen flash, etc.) in a progressive enhancements manner.
So Arc can't even generate fully functional links, let alone keep up with modern web advancements.
Good grief, that took a long time. Here you go: http://pastebin.com/bSW5dfRQ [1]. I'd better stop neglecting my duties now!
Edit: One thing I forgot to put in there: one reason the closure technique is powerful is that you're leveraging the programming language and runtime to do most of the book-keeping for you. Whatever data is handy, you just reference. The system will remember all the references. That's why using things like query strings and hidden form fields is more complicated: you have to handle all those details yourself (not to mention serialize and deserialize them if you're passing through any other format than what your program keeps in memory). That is tedious, and when your app has many kinds of request, the complexity quickly piles up.
Of course there are other abstractions you can build over this, but closures are an elegant one—especially in cases where programming simplicity is more important than scalability, which is most cases.
Edit 2: A few people thought this should be its own post, so I made https://news.ycombinator.com/item?id=8425011.
[1] Originally http://pastebin.com/dETyYtpX, but I added the above bit etc.
Best news I've heard all day!
quick fiddle to dump the json
Or should I be picking them up from "/v0/topstories"?
This should really energize the HN community.
It will be nice to write that directly via the API :)
have a name ? I can't actually find anything useful on the 'net.
I feel like I must be missing something... :/
If you're just planning on displaying the text in a browser, no decoding is needed. If you want to parse the text to do some sort of textual analysis, an HTML parser library might be best.
You've given me an idea though, so back to vi for me.
Thx.
http://us1.php.net/manual/en/function.get-html-translation-t...
absent another source, you could dump it out for your usage elsewhere.
% php -r 'print_r(get_html_translation_table(HTML_ENTITIES, ENT_QUOTES|ENT_HTML5));'
or % php -r 'print json_encode(get_html_translation_table(HTML_ENTITIES, ENT_QUOTES|ENT_HTML5));' | jq .
edit: just found http://dev.w3.org/html5/html-author/charref (but might be harder to parse..)Any idea?
Why is so much preparation necessary to redesign like 3 simple templates. Shouldn't it take like ~10 hours for one person to do this? I'm talking about the front end.
Why wasn't hacker news optimized for mobile a long time ago?
:|
EDIT: Also, the "about" value in the users api appears to be truncated. Compare...
https://hacker-news.firebaseio.com/v0/user/angersock#
...with...
https://news.ycombinator.com/user?id=angersock
EDIT2: Note that the JSON is correct, but the preview in the firebase API seems to be broken.
EDIT3: No issue tracker on the Github? Laaaame.
What you're asking for is a string you have to parse. That's a lot more work and there is a lot more that can go wrong.
[0]https://stackoverflow.com/questions/249760/how-to-convert-un...
var time = new Date("<ISO-8601 string>") // in Javascript
time = DateTime.iso8601("<ISO-8601 string>") # in Ruby
Aaaand it's human readable! Aaaaand it works before the Unix epoch!EDIT:
Changed DateTime.parse to DateTime.iso8601, to be even more retentive.
Look, parent claimed that parsing ISO strings is hard (it's not, especially if you're consuming a web api on a modern web language) and that it was more readable (which is so clearly wrong I have no words).
As for being more accurate, again no. The range is worse (lol wraparound if you're using a 32-bit int), there isn't explicit support for fractional seconds, doesn't map onto UTC cleanly, doesn't handle leap seconds, and so on.
It's only "easier" if you don't actually care about a human-readable timestamp that is robust and if you desire to do date parsing yourself instead of using any of the well-established libraries out there. Ugh.
Human-readable is a slight advantage, agreed.
Why not, let's have some more examples:
In PHP (with a bug, because PHP is stupid and doesn't handle decimal fractions):
$dateTime = DateTime::createFromFormat( DateTime::ISO8601, '2009-04-16T12:07:23Z');
In C#: DateTime time = DateTime.Parse("2010-08-20T15:00:00Z", null, DateTimeStyles.RoundtripKind);
In Python (stupid that it isn't in the standard library, see https://wiki.python.org/moin/WorkingWithTime): import dateutil.parser
dateutil.parser.parse('2008-08-09T18:39:22Z')
In Perl:
my $dt = DateTime::Format::ISO8601->parse_datetime( "2008-08-09T18:39:22Z" );~
Interchange formats, say JSON blobs over a wire, should very clearly express what's in them, perhaps by using a very well-known standard which is human-readable, whenever possible. The fact that some languages haven't yet realized that this is an important-enough feature to put in their standard libraries compared with whatever esoteric academic shit they think is necessary (C++1x, for example!) is not the format's problem.
Hint: if you're consuming a web API, you are probably using one of the languages I gave examples for, or a very close relative. Just because your Haskell-on-M68k package doesn't know how to 8601 doesn't mean that using a nondescript number is a good idea.
Why, in the year of our Lord 2014, is this even a fucking question?
EDIT:
I'm sorry to be so mean in my language about this, but I've had to fight a lot of raging stupid with regards to storing timestamp data. I do not wish to see anyone else suffer unduly.
Funny you should say that, UNIX time works great for log files. Since now even a dumb tool can search for values in the range, 1388491199-1420027199 to see all 2013-2014 lines.
Your way requires specific support for the date format, dealing with "2014" false positives, or searching through all 12 months individually (2013-01, 2013-02, 2013-03, etc).
You keep drumming on about library support, which is important for a format which isn't universally supported. Fortunately for UNIX time there's no value in listing them off one by one, you can just assume it is all of them...
UNIX timestamps are also listed in several international standards. Including POSIX.
Again, no benefit to not using it, other than you are scared of strings. If you are handling JSON from a web API, you are already dealing in strings.
I never claimed that. In fact I didn't address readability at all. So I have no words for your "no words" relating to a claim that literally didn't appear at all.
> The range is worse (lol wraparound if you're using a 32-bit int), there isn't explicit support for fractional seconds, doesn't map onto UTC cleanly, doesn't handle leap seconds, and so on.
Fortunately we're already well on our way into a 64 bit world, and aside from legacy systems it won't be a problem by 2038. According to the Steam hardware survey [0] over 80% of Windows machines, 100% of OS X machines, and 90% of Linux machines are already running a 64 bit OS.
Leap seconds can be handled during the cultural conversion.
> It's only "easier" if you don't actually care about a human-readable timestamp that is robust and if you desire to do date parsing yourself instead of using any of the well-established libraries out there. Ugh.
I don't care about human readable timestamps for an API used in automation. More robust is subjective, particularly as parsing it is more technically complex (particularly as most of the parsers support several different but similar DateTime formats).
Most well-established libraries support UNIX time natively or use it internally.
There are very few good reasons to choose a Unix timestamp to represent a date when compared with ISO-8601.
> There are very few good reasons to choose a Unix timestamp to represent a date when compared with ISO-8601.
Complete lack of string parsing is a good one.
Examples are coming soon, but we didn't want to make people wait.
There's also certainly a big support overhead dealing with the moderation of stories and comments.
I'm not saying it should take hundreds of dedicated engineers to run the site, but I think it's silly to look at it and think you could run it in your spare time because it's just a list of stories.
I always assumed HN was more of an operations challenge rather then a purely algorithmic one. Lotsa traffic I assume. Not sure about the scope of user/comment moderation around here but something tells me their item ranking system just might be that good.