That's the kind of facile statement that makes people riotously mock the entire startup community, like "MongoDB is webscale" but even less valid.
Cloud services are not a panacea, and there are myriad situations in which running one's own infrastructure can be a good idea. What matters is that the issues and benefits are taken into account; if one can show research demonstrating that a custom infrastructure is cheaper, or more reliable, or less prone to legal issues, for example, then there's nothing to laugh at.
And remember that PaaS in particular can cost a buttload of money - I'm certain it's contributed to the downfall of more than one otherwise promising startup.
[1]http://blogs.computerworld.com/18198/oops_amazon_web_service...
Do you even have access to that snapshot if the system is down?
the cloud simply exchanges know levels of failure with unknown level of failure. You pay someone else to think about it for you.
If there is a massive data loss at amazon, and people have been backing up to glacier, the recovery times will jump from hours to days even weeks, simply because there aren't enough drives to recover from.
The cloud may be more reliable, but when it goes down, it'll gown down hard, leaving you high and dry if you don't have real backups.
Besides, HN also uses cloudflare as CDN and DDOS protection, so it's not like they're stuck in 1999.
This is quite hard for startups to do, because their core function is to keep adding and removing features, experimenting and scaling, and the initial cost of buying hardware to support these functions is too large. This plus all the services provided by AWS et al. saves a lot of time and effort.
Nothing to do with "cloud" vs "non-cloud".
Repeated Disk corruption maybe indicates a hardware problem unless the file system isn't up to snuff. Generally those file systems are pretty reliable.
We really liked AWS, and the support we got for most of the times was good, but as we had 24x7 live traffic, even with reserved instances, the cost finally caught up to justify the move to our own hosting solution.
One thing that the made the migration easier for us was that from day 1, we decided to treat AWS as a co-location, hence we set things up with the usual open-source s/w stack (i.e. avoided proprietary Amazon solutions like dynamo, messaging solutions etc.. maybe just used S3 for offline archiving of logs. It was tempting to build on their components, but ended up building it ourselves.), including own own hadoop setup. When the time came, we could easily migrate things out.
Hope this gives some perspective. Still a fan of AWS, and if I were do another start-up, would follow the same script all over again.
Also, I suspect that if a startup came to them with a loyal following like HN enjoys, they wouldn't be laughed out of the room at all...
This doesn't explain the penny-wise (IMO) allocation of hardware or dev/sysadmin resources of course.
Dropbox was an rsync executed every two minutes.[1] AngelList was done via email.
[1] can't find this reference atm, could be apocryphal
It seems to incur maintenance costs and reputation hits that are avoidable though, isn't that reason enough to invest some time in building a reliable solution?
Besides everyone knows, nobody whistles 2600Hz, they just get the toy out of the Cap'n Crunch box to do it for them.
Kidding aside, running a bare metal server you own or rent is always cheaper, assuming you cannot save money by turning things on and off and know you capacity needs. Sure HN grows but not nearly as fast as FB, Twitter, etc. the service is big enough where it would require expensive virtual servers. This is the best case scenario for a physical hardware box run by people who are familiar with such things.
They even went out of their way to get the fastest single core xeon possible, which is a fairly mid-range CPU, singe they care about process performance vs. overall.
Yes it may seem archaic, but it a real server has real disk IO, something which on amazon and the like doesn't come cheap
In many circles, you would be laughed out of the room for cargo culting like this, and for trying to draw a bias against a legitimate deployment choice by declaring it "legacy".
There are many scenarios where hosting on your own hardware is a superior option for a variety of reasons: Financially, security, performance, flexibility.
In the case of HN it seems like it's on some pretty meager hardware, making compromises like software RAID. If this were a critical system for YC, they would have it on redundant machines with redundant, flash-based, hardware-RAID equipped platforms, clustered with redundant 10Gb cross connects, etc. Criticizing a deployment strategy because of the peculiar issues they have faced is like writing off AWS because someone's unbacked up small instance got killed and they had no strategy for it.