A gentle reminder that RAID doesn't make offsite backups
journalspace.com
journalspace.com
RAID isn't a backup solution, period.
http://web.archive.org/web/*/http://journalspace.com
"We're sorry, access to http://journalspace.com has been blocked by the site owner via robots.txt."
Otherwise, a tool like 'Warrick' might get back a significant portion of the content -- especially any that was old and well-inlinked. As it stands, maybe they can scrape a little from Google's cache.
Warrick: http://warrick.cs.odu.edu/
RAID isn't a guard against filesystem bugs, carelessness, or extreme physical server damage, but it seems to me this is more of a warning against bad blood.
Somebody has to eat the broccoli and do the shit work of making sure that the data is backed up; if you can't do that or you hand off the job to the sort of person who is going to sabotage your company because they don't feel appreciated... you aren't a startup, you're a hobby.
On the other hand, it probably isn't a coincidence that this happened to someone who had no backups at all.
In this case that could mean having 3 USB disks, each with their own 'owner'. The owner of each disk would be responsible for making the backups and storing them in a place only they can access.
This is also lesson on how to let someone go. They say it's possible that a former employee may have done the deed. When letting someone go on bad (and even good) terms you need to have IT disable all of their accounts as soon as possible, preferably while they are in-process of being fired.
That said, if this kind of problem is a valid concern for an installation, the simplest solution is external validation of backups. For example, one might hire someone not in control of the original system (and not in contact with them) to do restorations and provide certification that they worked according to a test plan. This is the foundation for a lot of Sarbanes-Oxley technical audit consulting.
Audited backups would, of course, require a fairly elaborate disaster recovery plan, which sounds out of scope for this company.
I don't want to pass judgement, but if a pivotal IT person was let go and sabotage was discovered, at least a basic pass through the systems taking backups doesn't seem out of line.
Pretty amazing.
That said, it's really hard to do backups of lots of data.
And anyway, if they can deal with rebuilding a RAID after a failed drive then (by definition) the site can deal with copying all the data off the drive. Heck, you could periodically yank a drive our of the RAID and replace it with a fresh one, and you'd have a backup.
Database backups almost universally have to be made by the database system. This is no excuse for lacking a backup system; backing up databases is a solved problem.
1) The app may be in the middle of a series of DB commands that all need to complete with success before the DB is consistent at the app layer.
2) The DB is in the process of writing out some table rows and hasn't finished.
3) The DB has written some temporary locks to portions of the database that need to be released.
4) The OS hasn't committed writes from the DB to disk
5) The disk hasn't committed writes from the os to platter yet.
Your best bet of this working is to cleanly shut down your app, then cleanly shut down your DB and run an fs sync. After all that it might be ok to yank the drive.
Just yanking a RAID1 drive may work sometimes, but I wouldn't count on it. Especially when every DB system I know of has some sort of backup/dump mechanism. As someone else mentioned, RAID is great for providing high availability, but it does not provide disaster recovery.
It backs up your stuff to S3, and it can work on your Mac/PC, or on Windows Home Server. Have't tied it yet, but looks pretty neat.
Also, new NAS from HP has built-in S3 backup: http://www.engadget.com/2008/12/29/hp-mediasmart-server-ex48...
I mean, we assume Google has a really amazing backup strategy for Gmail. We hope they do, because, jesus, I have a lot of irreplaceable information in there. But we don't actually know.
Hmmm... does Gmail have an "export all" feature...?
My operation is way smaller, and I have a cron script that runs once a week to dump the database, zip it, and then transfer it to a backup server.
This is one the first things that I set up. It really should be the first thing any company with user-generated content to do.
- case fans fail, everything overheats
- room AC fails, everything overheats
- power supply goes haywire, toasts drive electronics
- box falls off the shelf, drives crash
- fire or smoke damages drives
- roof leaks, drips onto drives
- one drive fails but nobody notices for a month until the next drive fails
- burglars steal box
No, not all these have happened to me.
The problem seems obscure...what kind of software bug overwrites all data on disk?
dd if=/dev/random of=/dev/sda1