They're far more expensive than e.g. Rackspace there for some reason.
They're far more expensive than e.g. Rackspace there for some reason.
Whilst this might not sound like an enormous issue, DRAM errors are surprisingly common in real deployments[2] and can result in scary things happening to your data. If you're using a machine without ECC memory make sure you're prepared to deal with any possible issues that might arise, especially if it will impact your core business.
[1]: http://www.hetzner.de/en/hosting/produkte_rootserver/ex8
[2] DRAM Errors in the Wild: A Large-Scale Field Study: http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
See this Stanford ee380 talk: http://stanford-online.stanford.edu/courses/ee380/100922-ee3...
The part about DRAM errors is about 57 minutes in.
Abstract here: http://www.stanford.edu/class/ee380/Abstracts/100922.html
Also, AFAIK most hosting providers don't even have ECC ram as an option for servers, e.g. Amazon.
http://aws.typepad.com/aws/2010/11/new-ec2-instance-type-the...
Given that they list both the standard memory and GPU memory next to each other, but only put ECC next to the GPU memory, it seems relatively likely to me that the standard server memory is not ECC.
On the other hand, the fact that https://forums.aws.amazon.com/message.jspa?messageID=203167 has never been answered suggests they do not, as they would answer affirmatively if it was true surely. Of course they may use a mixture.
EDIT: Interesting that James Hamilton is on their team and thinks they should use it http://perspectives.mvdirona.com/2009/10/07/YouReallyDONeedE...
The costs involved were a lot higher 3-5 years ago than they are today.
Lots of things can happen, you should have a higher level of replication such that you can handle a whole server going poof, not just a single bit going poof.
The cost of ECC at Hetzner-- the cheapest provider out there- is about half an additional server. So, buy three servers without ECC for the price of two servers with ECC, and replicate your data three times (and triple your bandwidth, horsepower, etc.)
This is not hard with platforms like Riak which are distributed homogenous clusters of nodes.
And if your service isn't built like that, then really it should be. (IMNSHO, of course.)
Whenever data is read, you read from more than one replica, and then compare them. If one of them has been corrupted, its hash won't match and you'll know it. You can then write out the correct data to the node with the error. This is very easy in systems like Riak.