Postmortem of last week's fileserver failure
github.com
github.com
The problem with the GitHub RAID approach + daily snapshot is that you end up with only the equivalent of a daily backup and you do not have data integrity insurance.
Each of their file storage pair is running in RAID + RAID over the network (DRBD). RAID is not a backup because you can get data corruption and you just replicate the corruption.
Imagine that they got part of their A server having issues, which are with current FS+RAID simply ignored. You can run months without integrity check even with Git because if you add new objects and do not need to repack, no integrity check is performed on the old objects. So your corruption spread from A to B (RAID) to backup (daily snapshot). Then you need at home to checkout because you lost your HD, you are doomed. Full data loss without a single warning. If you are a single developer on your own private project, this scenario is easy to achieve.
I really really like the way Fastmail is doing backup. They are replaying the IMAP operations with checksum. With Git you can also replay with checksum too. In fact, it is built in thanks to the DSCM nature of Git.
For the backup of my customer git repositories, this is what I am doing, in the post-update hook, I fire a git sync with git on another server, it is doing everything needed while being CPU+bandwidth efficient. Thank you git.
The easy way to do it: http://kerneltrap.org/mailarchive/git/2007/10/18/346839
Presumably the anomalous behavior was disk-related and yet they only test memory.
Why not check SMART status? Or if the RAID driver doesn't support SMART passthrough; move the drives to the SATA controller and boot off of a flash drive.
Also, there are tools to generate lots of i/o and test for failures, i.e.: