Cause of YC/HN outage discovered
archub.org
archub.org
Interesting how easy it is to move your whole web site. Good customer service is important when users can switch so easily.
Hopefully that's the normal of months ago when loading comments/threads etc was quick?
Still taking 10-20 seconds to load most pages. One thing I haven't tried yet is just creating a fresh account - maybe mine just has too much associated with it (Maybe I comment too much etc).
edit: ah I see it only affected static content. Shame :/
Edit: Slicehost help articles & resources, well organised & exhaustive is a plus though.
I'd like to highlight this lesson in how to loose (or not gain) customers by randomly shutting down technology sites that serve the decision makers you wish to influence.
Tom
P.S. Well done Slicehost.
When the traffic starts to impair neighboring sites, something has to be done. Just about any ISP will do the same thing: block the site with the surge, that could possibly make other arrangements, rather than inconvenience other customers whose traffic is as expected/usual.
The detail missing so far is why Pair noticed today, if it was the same level of traffic as before, or a slow build. Was a new threshold crossed? (Did someone's HN-focused tool go haywire?)
The Pair message suggests end-of-day logs will be the way to tell for sure.
As you say I doubt today was unusual for HN, it's the way Pair went from zero to shut-down.
One of the things I like about the way that Joyent operate as a cloud host is that they allow you burst on shared boxes because of those times you need it. At the same time they'll let you know you need to think about buying more resources without just slamming on the brakes.
Pair should do a much better job of noticing a soft limit earlier, so if a heavy traffic day had hit HN PG et al should have been already aware they were overusing their shared hosting and were planning a route out.
It currently is not really possible to track that per user in a shared environment and see who is using what resources from an individual perspective of MySQL / Apache / CPU / Mem.
Ofcourse it is.
Or how do you think they determined that it was pg's site causing the trouble?
I doubt it is a HN-Focused tool, as this affected the static content from www.ycombinator.com, and images and CSS are not the focus of bots usually.
Assuming it was not a DoS attack, a smart host should have noticed the traffic load increasing over time, and offered to upgrade to a less-loaded server and recommend a dedicated server.
By disabling the site, they have lost a customer, and lost on a up-sell to a dedicated server.
a) anything
b) except
c) killing their service
Of course, the flip side is that leaving it running adopts an attitude of "screw all our other customers, they can eat crappy service while we kiss up to the popular guys who are chewing up everybody else's server resources". Which isn't what I'd look for in a hosting provider...
Dear [account_contact_user]
Your website traffic has risen beyond the maximum threshold of [threshold_amt] for the [name_of_level] level of service.
Since we appreciate your business of the last [length_of_service], we have given you a 24 hour courtesy upgrade to our next level of service -- [name_of_next_level]. If, by [end_time] you decide to keep this level of service you must contact our sales center to arrange payment. Otherwise we will have to start throttling traffic to your server so that it remains below the threshold of [threshold_amt] and does not impact our other customers.
If you have any questions about this courtesy upgrade, or wish to keep this new level of service, please contact [account_manager] at [account_manager_details].
Thank you for using Pair Networks for your hosting needs.
I would have been happy to respond to an email saying I should upgrade & pay more, but even after a few emails with tech support that option didn't come up.
The only limit they advertise for that class of account is data transfer, and because it's mostly just serving uparrow.gif -type files we're well below that limit.
I don't get shared hosting at all. VPSs are dirt cheap these days.
More important to me is that I am really really really surprised that Ycombinator was running on a shared account. This just blew my mind.
I have no hard feelings towards pear, I would have shut the site down too; possibly long ago.
Please turn in your geek card at the door.
In short, Pair flaked, but we had in fact planned the system in a way that protected us against it.
It's nice for your hosting all of your little sites and getting them up in basically 0 time, even if you need wordpress or something. The panel's pretty nice. I'd at least sign up when they have one of their crazy deals (I got a year of hosting with a free domain for $9.something, which basically means you pay for the domain and get the hosting for free. It's a good way to give them a try with the main cost being that they'll suck you in and get you paying $9/month after your cheap price expires.)
I jumped on a deal at the beginning of this year for ~$9 for 1 year of shared hosting. I now use that account purely for running things like scripts through SSH for doing basic data grabs with wget for further analysis elsewhere.
I've more than got my money's worth and I've yet to see any kind of email message asking me to upgrade.
I moved to slicehost and appengine[disclaimer,etc] and have been relatively happy with them.
Yeah, it sucks that one of our favorite tech news sites was impacted by this, but how impacted were all those other customers?
It is easy to make a smartaleck comment about how Pair was trying to upsell by doing this, which is preposterous. Pair is a well-respected provider with many more years providing good service at a fair price than HN has existed, and I'd be willing to bet will be around after HN has peaked and begins to move back to the traffic load that might make sense on a shared system.
But the fault here ultimately lies with the folks running HN who thought it was wise or appropriate to host any of its content on a shared server that likely cost them less money per month than most of us spend on soft drinks in a week.
Fine, nobody's saying HN/YC should be allowed to overuse resources. I didn't say anybody was saying that.
Not giving a warning probably doesn't count as great customer service, but then again, once the problem had been identified by Pair, and once they knew of the negative impact HN was having on every other paying customer on that server, what kind of customer service to those other customers would it have been for Pair to fire off an email to HN then wait an hour, or thirty minutes, or ten minutes, before shutting it down?
How long should Pair have allowed HN to impact other customers to satisfy folks here? And what makes HN more important than any other paying customer on that server?
Oh right, it's because you read and like HN, which, ironically, so do I.
As for it being a sales screwup, maybe. I kinda doubt there is a great deal of overlap between HN readership and the average potential Pair customer. We could also suggest that Pair taking action to protect all those other customers on the server is an example of how they would provide good service to the many when they're being hammered by one overpowering fellow customer.
I think you should upgrade to a more "dedicated provider" rather than a "dedicated server"...
But seriously, not even a warning?
That is why shared hosting is cheap, you start with it and once you are successful or starting to get slashdotted you buy something bigger that can scale.
I've been in the industry for 10 years and worked for quite a number of hosting companies, not Pair though, and when you have 150 shared clients on a machine and 1 client is causing the problems you do your best to deliver the warning before it gets out of hand but it is very hard to do.
My server literally had its plug pulled.
Geez, and I thought that watching the janitors while the work was over the top, now I gotta watch the rumba too :(
What they should have done is upgrade the website to a dedicated server for free and let that news hit the front page.
Seriously-they couldn't just email the account holder?
And of course they are going to disable it without warning if it's causing problems for all the other customers on that box...
There was no sudden spike in traffic. If they'd bothered to check the logs they'd have found that the load, whatever it was, was no higher than it had been.
It makes it easy to minify, combine, gzip and push your css, js and image assets into the Amazon Cloudfront CDN with far-future expiration headers. It also automagically detects background images referenced in your css, and puts them in the CDN. It rewrites the css to use the new CDN image urls.
It's all done through a trivial REST API.
Would love some feedback, and to find out if/how it's breaking any of your complex css/js.
Dad used to get a shared hosting account for things like that, but now we have services that host content for free. They are much easier for Dad to use as well.
In this age of $20/mo VPSs and free content hosting, I'm honestly not sure how shared hosting survives. It's not as flexible as a VPS for hosting sites (and with more than two sites, isn't even as cheap), and it's not as easy to use (or as free) as Facebook.
John Smith Plumbing's five page brochureware site doesn't need the headaches of having a VPS but also won't accomplish business needs on Flickr.
VPS or cloud is dirt cheap these days, no excuses for not planning growth?