HN is not a service, its not a startup trying to get mindshare - its a (luvvie alert) community - and one where the downtime had a known and obvious cause and did not matter - rtm would fix it, I would catch up, I knew this thread would exist and I would still get insight into what is happening 12 hours ahead of everyone else.
I thought I would be more bothered - turns out I am emotionally invested, not emotionally dependant.
To PG et al, good luck with the fixes, and thanks.
Cheers
From my perspective this looks like either a huge misunderstanding or the definition of straw man rhetorics. And I have to say I don't get where your hostility comes from, but I neither recognize my message as it was apparently understood by you nor do I appreciate the tone. I'm always ready to admit that issues like these could be my fault as I all-too-often like to hide behind the fact that I'm not a native speaker. I just don't see it this time though.
In the spirit of being constructive: I don't think the application of the label "normalcy" or the lack thereof is a value judgement in any way, and being "normal" is something that must be interpreted within the given context - otherwise it makes no sense because it's not a global property that encompasses the entire personality of a person.
Regarding the actual subject of normalcy I'm finding it difficult to add anything else without repeating myself. From personal experience, there is always "good normal" and "bad normal", as there is "good not normal" and "bad not normal", but even more often it just means "different from the other 80%". In this case, I chose the word normal because it instantly conveys the fact that most hackers use the web in different ways from most of the other people. That's all. It did its job.
It's unfortunate that this triggered something in you, but I encourage you to take a step back and recognize that you saw some things in my message that were simply not there.
I recognize it might well be an expression of a cultural divide between us, but this is the way I see it: HN is populated by a huge group of very diverse people. While we are diverse in pretty much every aspect, we are also bound together by being hackers and by living, at least party, in whatever hacker culture means for each us. That's not normal, it's a subculture. Almost every single human being is part of at least one subculture, and many are members of several, and there are millions of overlapping points that defy categorization. I cannot for the life of me figure out why it would be considered wrong to express this fact.
I can say for sure that I am not normal in many aspects, and then again I am normal in a lot of others. It's not somehow evil to recognize and talk about this. I think before dfc's comment, the number of people who misunderstood my original comment was very low, but my point is tainted now, subverted and burnt beyond recognition.
"just one summer with the RSC back when I was younger".
"Of course Johnny said, ... umm, Gielguld of course dahling, did you see him in the Tempest in 74? Wunderful"
"I was saying to Sam just after that Quantum thing. You really should talk to Barbara about directing the next one. I have her number; charming lady, did I ever tell you about the time she ..."
A small red rotating flashing light will rise up out of the coffee table next to you and the words "luvvie alert" will be flashed across your field of vision.
EDIT: More proof -- downvoted -- what could be more hellish than to ban humor?
BUT I'd much rather get rid of the "this link has expired" syndrome.
Fast is only impressive if the results are modern.
I do that much more often than I hit an expired link.
Here's an example. Currently if you go to the home page and then click the "More" link at the bottom to go to the next page, and wait long enough, the more link expires. That's because it's currently implemented as something like this (not sure about ARC syntax, I'll use Scheme syntax here, but you get the idea):
(define (show-list-of-posts page-number)
... display the rest of the homepage ...
(link "More" (lambda () (show-list-of-posts (+ 1 page-number)))))
This stores the closure `(lambda () (show-list-of-posts (+ 1 page-number)))` with the free variable `page-number` in a hash table on the server. The URL becomes an index into that hash table. But the only information needed to reconstruct such a closure, is the function body and the free variables. So if we defined an auxilary function: (define (foo page-number)
(show-list-of-posts (+ 1 page-number)))
then we could represent the closure as the pair ("foo",page-number). If we encode that in the URL, then instead of looking up the closure from the hash table, we can reconstruct the closure on the fly. Hence we do no longer have to store anything on the server, and no links can expire anymore.There are some challenges when you want to do this transformation automatically, but they can be overcome.
I don't think this is strictly true. The comments and reply links simply link to a normal (non-closure) URL on the site and the page is generated from that URL in the normal manner.
The reply links have a parameter 'id' - the id of the comment to which the reply is being posted. I would guess it's just passed to a normal function which adds the reply to the post.
May be you mean the same thing, but "serializing closures in the url"(where? all I see is the id of the post being replied to) isn't the same as params in the url which are then passed to a function.
If you have a lambda expression like `(lambda () (show-list-of-posts (+ 1 page-number)))`, the Scheme implementation does these things:
1. Introduce a global function definition with the same body as the lambda expression but with an extra parameter for the free variables. The free variables in the body of the function get replaced with expressions that extract their value from the extra parameter that has the free variables.
2. Convert the lambda expression to an expression that builds a pair of a function pointer to that global function, plus the values of the free variables.
For example the code:
(define (show-list-of-posts page-number)
... display the rest of the homepage ...
(link "More" (lambda () (show-list-of-posts (+ 1 page-number))))
Will get converted to something like this: ;; this is the extra global function that has the body of the lambda expression
;; note that the reference to `page-number` got
;; replaced by `(extract-value "page-number" params)`
(define (closure-324 params)
(show-list-of-posts (+ 1 (extract-value "page-number" params))))
;; note that the lambda expression gets replaced by a create-closure expression
(define (show-list-of-posts page-number)
... display the rest of the homepage ...
(link "More" (create-closure closure-324 "page-number" page-number)))
create-closure creates a closure data structure where the first argument is the function pointer, and the rest of the arguments are the free variables.Now, what happens if you want to remove this closure business in a web application, and instead use normal URLs?
First, you introduce a global request handler for the body of the lambda:
(define-handler (post-list-handler params)
(show-list-of-posts (+ 1 (extract-value "page-number" params))))
This would define the handler for news.ycombinator.com/post-list-handler?page-number=12.Then, instead of the (link ...) with a closure, you just link to that url in the show-list-of-posts function:
(define (show-list-of-posts page-number)
... display the rest of the homepage ...
(link "More" (create-url post-list-handler "page-number" page-number)))
Compare these code snippets to the one above. Do you see the similarity to closure conversion? In both cases we:1. Introduce a global function/handler for the body of the lambda.
2. That function/handler gets a `params` argument that has the free variables.
3. Everywhere a free variable is referenced in the body, it gets replaced by an expression that extracts the value from the params argument.
4. In place of the lambda expression, we have respectively a (create-closure func free-vars...) or a (create-url handler free-vars...)
So it's really completely analogous. That's why I say that we are just serializing the closure here, and this could be done automatically. Hopefully this makes it more clear what I mean, but maybe these details just make it less clear if you're not familiar with how closures are implemented (closure conversion)...
Great explanation of what's causing that - thank you!
I suspect hindering spam and vote manipulation plays a part in the architectural decisions. It also makes it possible to create different HN's for different users - e.g. the hellbanned, royals, and plebeians.
This is particularly infuriating on a cellphone.
One quick fix: in the code that prints this message, check to see if the request was a comment-post action. If so, append, "...but for your convenience, here's the text you tried to post, so it's not lost forever: [...]"
(Just watch out for XSS!)
After seeing this thing down for several hours -- and Google hate downtimes, they punish you instantly -- I think: "Maybe not."
However, happy to hear what went wrong and why we still should go for the bare metal thing.
seems the issue was more with the migration.
so i reckon you're good to get back to removing all the threads from your apps :)
EDIT: since Heroku fooled Rapgenius and us all it would be one more reason to get into system operations again and host on bare metal.
Timing some pages, the response is around 350ms. That means 95% of the loading time is networking, not generation time. You're right, the server is really fast given the load HN brings down on one server!
Then to utilize this you want a load balancer in front that's unlikely to fail and if it fails can be restarted rapidly. This load balancer should send requests to the other server as soon as one has failed. The other option is DNS failover, which can work but has different trade-offs.
One thing that this does not protect you against is repeatable failure due to a software bug. If a request comes in that happens to crash one of your servers, load will be quickly switched to the other. But the user that cause the original server to crash might refresh his page because he didn't get any response back, which also sends the problematic request to the other server, and might crash that one too. These kind of problems are very hard to deal with (if the requests don't come from humans but from programs that automatically retry their request to other servers when they don't get a response it's even worse: a single bad request can bring an entire cluster down in seconds).
Paradoxically sometimes measures to improve availability can cause availability to go down. If you make your architecture more complicated you introduce more opportunities for failure, especially due to bugs. If you have multiple servers in a pipeline, for example a load balancer and then the real servers which in turn talk to the database servers, that can also cause your likelihood to fail due to hardware to increase. If your pipeline depth is three, then the chance that you have a hardware failure is about three times as big compared to when you had a pipeline depth of one. You want to minimize the depth and maximize the width of the pipeline.
So you should ask yourself whether it's worth it, especially when you're small. Many of the successful sites had or still have a lot of downtime. Maybe it's better to minimize the duration of downtime with fast restarts rather than trying to lower the probability of downtime.
I did visit one of those quickly and the result was HN should be reachable... Hmmmm, therefore I. Checked with another age and this also timed out on my tablet. So either my ADSL or the server was problematic? I went for dinner and upon return it worked.
The check I used was http://www.blockedinchina.net/ First time I used this site...
When I saw http://www.downforeveryoneorjustme.com/http://news.ycombinat... I realized that I wasn't.
But you could have easily met me. Mostly Linux events... mail me if you want to know
Apologies for the HN outage. We think it may be related
to the new server. This could take a while to fix.
[1] https://twitter.com/paulg/status/303350671460147200Just because the switch-over went smoothly doesn't mean that normal operations didn't fall apart later on.
Alternatively you know... you can just accept that the service will be down for an hour or two and most people wouldn't even notice without this post.
Bonus points if you do actually. Their search API is limited to 100 a day, which is not a lot.
seeing submissions in realtime, reducing load on the HN server. Saving F5 addicts.
Sorry, I don't know Lisp and Arc seems to be a dialect of it that I also don't know.
Overcapitalization is annoying, and it's "it's", not "its".