The ECC RAM absolutely does find and corrects problems (we see them in the logs). However, just to be absolutely clear we would not need ECC RAM -> Backblaze checksums EVERYTHING on an end-to-end basis (mostly we use SHA-1). This is so important I cannot stress this highly enough, each and every file and portion of file we store has our own checksum on the end, and we use this all over the place. For example, we pass over the data every week or so reading it, recalculating the checksums, and if a single bit has been thrown we heal it up either from our own copies of the data or ask the client to re-transmit that file or part of that file.
At the large amount of data we store, our checksums catch errors at EVERY level - RAM, hard drive, network transmission, everywhere. I suppose consumers just do not notice when a single bit in one of their JPEG photos has been flipped -> one pixel gets every so slightly more red or something. Only one photo changes out of their collection of thousands. But at our crazy numbers of files stored we see it (and fix it) daily.
Have you tried using ZFS on one of these?
ANOTHER option down this line of thinking is switching to btrfs, but we haven't played with it yet.
By the way, at Backblaze I've felt like we have had to implement several things that I would have guessed would be standard "off the shelf". One example: When a customer wants to download a restore file with a web browser, are you aware there are no checksums for over-the-network transfers other than the built in (completely unacceptable) 16-bit TCP checksum? You are virtually guaranteed to have an undetected corruption within 30 Gbytes of download, which is basically what we like to call "a totally average customer restore". So Backblaze had to write our own custom reliable, restartable "downloader". It boggles my mind that the whole internet is throwing undetected errors on HTTP downloads and nobody cares to fix the protocol?!! Where the heck is Google, Apple, Facebook, Microsoft, or <insert standards body> defining a standard for web browser downloads larger than a few GBytes?
As for HTTP, I agree with you entirely! There is actually a standard that lets you put a hash (SHA-1 or SHA-256 or something) on an http anchor link, and the browser will verify it when the download finishes. It hasn't really gained very wide adoption though. Personally, I think something akin to Bittorrent is a better solution, since it doesn't have to redownload the entire file when it detects an error. It's ironic that our videos often have better data integrity than the graphics driver installer that I just downloaded, or the browser that I downloaded yesterday.
Two possible ways of solving this: make a download into e.g. 16 MB chunks and append a SHA-1 checksum for each chunk (not enterely unlike BitTorrent). Then re-download the chunks where the SHA-1 doesn't match. Another solution would be to use e.g. Red Solomon error correction.
Microsoft does the same thing -> they have a custom application to download their OS updates reliably. They do not use Internet Explorer's regular download capability because of the limitations.
It baffles me that the big browser companies (Google with Chrome, Apple with Safari, Microsoft with IE) are not interested in building a standard reliable downloader. I mean, what else does a web browser really do?
I'm not sure which HTTP client implementations reliably test those checksums, though. (Or whether CRC32/Adler32 would be adequate for 30 gigs!)
But brianwski's point remains that neither HTTP encodings, nor HTTPS can detect certain errors like in-memory corruption of the data after it has been read from disk, but before it has been sent out over the HTTP connection.
The last time we tried Linux for this, it couldn't do it. Maybe that's gotten better though.
We also tried DragonflyBSD, but it had hardware support issues, and OpenBSD, but unfortunately OpenBSD just simply cannot do large filesystems. At all.
ZFS-on-FreeBSD is the way to go at the moment, I think.
I actually use ZFS on Mac OS, as well, and have been happy with it there (thanks Dustin Sallings!).
But the Supermicro board says it can: http://www.supermicro.com/xeon_3400/Motherboard/X8SIL.cfm
The board also supports Xeon processors, so it might be that it only supports ECC with Xeon. However, I didn't see anything in the docs to support this, instead it says "Dual Core processors of Ci3 and Pentium: support ECC UDIMM only"
I'm curious since I ended up springing for Xeon in my system specifically for ECC. Now I'm wondering if I made a mistake...
"The X8SIL/X8SIL-F/X8SIL-V supports up to 16GB of DDR3 ECC UDIMM or up to 32GB of ECC DDR3 RDIMM (1333/1066/800 MHz in 4 DIMM slots.)"
Below are what we assume are ECC errors from an i3 based Backblaze pod, when this happens it is usually bad RAM (bad enough to crash the pods, replace with fresh RAM repairs it):
### IPMI LOG: ####
ipmitool sel elist | grep -v "Fan FAN" 69 | 01/04/2011 | 20:21:52 | Memory | Correctable ECC | Asserted 73 | 01/04/2011 | 21:15:50 | Memory | Correctable ECC | Asserted