Let's Stop Talking About "Backups"
joelonsoftware.com
joelonsoftware.com
The weird part is that you didn't mention the 400 pound gorilla in the room. ("I've been thinking about this for a long time, but it really hit home recently when...")
To my mind, that's why the post seems odd or passive aggressive. (It's a relatively short post as well. It just feels clipped somehow. The reader inevitably says, "What's not here?")
The web has more than enough content that feels stale after a week. Better to aim for something with a slightly longer shelf life.
But I do like his idea. So I appreciate that he omits pointless personal details and just sticks to the meat of the article.
When a writer expects me to read a bunch of self-indulgent cliche boilerplate I usually hit the back button.
No, because restores wouldn't have helped Jeff Atwood. You can backup a database onto the same server that hosts the database, and still successfully restore it later.
We should instead be talking about the concept of Continuity of Business. Decide how important it is for your data to continue to be available (constant availability vs. several days downtime vs. don't care if it disappears entirely?), then make a plan that gives you that availability.
What I'm describing isn't necessarily complicated. For example, I keep all my backup data (photos, music, etc) on a USB drive. Every morning it's rsynced to a hot spare. Every few weeks, I swap the spare with another one in a safe deposit box at the bank. (This means that we can recover from both immediate accidental deletions, and also ones that we don't catch for a few days, and disasters like theft or fire). I check every few days to make sure the drives are still working, and my wife knows how to switch the drives and has a key to the safe deposit box, just in case something happens to me.
Anyway, we feel that these are appropriate steps for protecting the data that we don't want to lose. But the important point is that everyone should make a plan that's appropriate for them.
"If you’re running a web service, you need to be able to show me that you can build a reasonably recent copy of the entire site, in a reasonable amount of time, on a new server or servers without ever accessing anything that was in the original data center."
We're tech types, but e.g. I know a programmer who never backed up his laptop (with important stuff) because he'd never had a hard drive go bad on him. You'll never guess what happened...
The hard part is getting 'normal' people to back up. It's really hard. They have all their data/documents on that laptop, and all the digital photographs from the first 18 months of their baby's life or whatever, but if you advise them to back up, you sound like some crazy lady who throws cats at people in the street.
that's probably why companies such as dropbox are doing so well. making backups/restores easy is sweet.
(ps - thanks superduper! http://www.shirt-pocket.com/SuperDuper/SuperDuperDescription... "heroic system recovery for mere mortals")
Time Machine is the solution: everyone gets it (at least the normal people I introduced it to).
So this is aimed one step farther up the chain.
Or, rather, it's aimed at all of us. It is a lesson that a lot of people need to learn.
That's just common sense. What happens if your hosting company decides to be dicks about some billing dispute and holds your data hostage? Even the if hosting company-powered backups work, you're still hosed without your own backups.
Simple backups are easy. Understanding when something is sufficiently backed up and how to know whether it is backed up can be quite complex, particularly if large amounts of data is involved.
USB drives work great for a desktop with 1 TB, but not so well for 20TB. Then there is the question of how you know whether everything is sufficiently backed up, whether it is free from corruption, how you will move the backup data to a live server in the event of a restore.
Remember: great people discuss ideas, normal people discuss events, shallow people discuss other people.
Most People who work in Production Operations environments of any scale, discover what has taken Joel the better part of a decade, in the first two-three years of their career.
I almost feel like that Airplane passenger sitting beside Brooks Jr - Brooks saw him reading his book, "Mythical Man Month" - and asked the guy (who had no idea who he was sitting beside) what he thought of the book - The gentleman responded that it was basically a summary of things he knew already. Joel is a giant in the industry, but he does have a tendency to discover/restate the obvious.
"It's not backups, but the restores that matter" - is kind of the mantra of every single person who has ever been responsible for backups.
Then, you go to _any_ class on running a production environment, and you discover things like RPO, RTO, Dress-Rehearsals, etc.. and the whole "It's restores that matter" begins to look quaint.
http://www.amazon.com/Practice-System-Network-Administration...
So this seems like a good lesson to take about backups. Mebbe three? One by your hosting provider, one at tarsnap, one on a separate dat tape, one on a usb stick?
A good point is though that even something as big as a dat tape looks pretty small by the standards of what we need to back up today.
Don't blindly rely on your partner to do it... Trust, but verify.
The IT company that I work for creates a backup system based on the requirements of our clients and then demonstrates the whole backup and restore procedure to make sure that it falls in line with what our client actually wants. It's really not difficult to do. Sure, some of the restore procedures may be slower (depending on other requirements, such as cost,) but the client knows that will be the case and signs off on it.
This is another reason why I like EC2 deployments: it is fairly easy to take your backups (automated deployment scripts, application, data) and spin up another copy of your whole system (except for flipping the DNS). Make sure those EBS-backed EC2 AMIs are really bootable and functioning :-)
I would think that if they are making money with AWS then they will keep buying more servers.
I usually trust S3 for backups and restores but periodically back my own data off of AWS to local storage.
I expect both Amazon and Google infrastructure services to experience outages from time to time. However, they have far more resources and expertise than I do to provide scalable services for a low cost.
Most places have very reliable backup procedures. Most of those have very poor restore procedures - I'd say about half fail when put to the test.
Since the restore is the important thing, that's the one you have to test. And if you haven't tested restoring, your backup is (quite possibly) worthless.
it's much better to ask yourself how long to replicate your existing system then how to back up. pxe boot to a kernel that you can install over the network with, bcfg2 to get the thing up to spec, start copying data.
a lot of machines can be back and configured in 5 minutes.
that said, i'm not you. i don't have terabytes of data to do statistics on. maybe there are other horrible details i'm forgetting. fast rebuilding is a pretty awesome strategy for a lot of cases.
Yes.
Why should I bother to write this? I'll outsource the task to the authors of High Performance MySQL, Second Edition, page 475:
Backup Myth #1: "I Use Replication As a Backup"
This is a mistake we see quite often. A replication slave is not a backup. Neither is a RAID array. To see why, consider this: will they help you get back all your data if you accidentally execute DROP DATABASE on your production database? RAID and replication don't pass even this simple test. Not only are they not backups, they're not a substitute for backups. Nothing but backups fill the need for backups.
--------
#!/bin/bash
HOME=
date=`date "+%Y-%m-%dT%H:%M:%S"`
rsync -aP --link-dest=$HOME/backup/current /home
$HOME/backup/back-$date
rsync -aP --link-dest=$HOME/backup/current /etc $HOME/backup/back-$date
rm $HOME/backup/current
ln -s back-$date $HOME/backup/current
#see if the disk is getting full
FREE=`df -lk|grep sdb1|awk -F" " '{print $5}'|awk -F"%" '{print $1}'`
#alert me if the backup disk is getting full.
T=80
if [ "$FREE" -gt "$T" ]
then
df -lk| mail $myaddress -s"disk alert $T% capacity"
else echo "backup disk is less than $T full"
fiI thought delayed replication was one of the main strategies they advocated in that book. I don't have it on hand. my mistake.
It's also not useful if your main files get corrupted and you diligently propagate the corruption. See: ma.gnolia
As far as the database goes, I'm big on stored procs + archive tables, but i'll leave that to the grown ups ;)
Whereas, a policy of automatically backing up everything except for your exclude list would have saved your bacon.
I don't mean to be rude, I don't know anything about you, but if you were a system administrator working for me, today would be your last day.
It's like reading an airline crash report, which I'm told often ends up sounding like a comedy of errors. Most airline crashes have an entire handful of causes, all of which are individually innocent, but on the one extremely rare occasion when they all happen at the same time they add up to disaster.
I'm always amazed at the number of people who spend their time thinking about replication strategies without understanding that data can be accidentally deleted in production too. I guess backups look "easy", so it's not as sexy an area of architecture planning.
Job security note for sysadmins: When someone suggests a disaster scenario, don't open your response with a laugh.
LOL! And if you get a corrupt block on your primary site, what're you going to do? All your standbys are instantly tainted!
Better leave this one to the grownups.
Spinning disks are good though. A lovely spot for backups. Just put them in a different building.
Edit: In the time it took me to write this 5 other people also lambasted this poor fellow. Ouch.
Redundancy is not a backup. If someone with full admin control to your system can destroy all of your data then you do not have backups. A proper backup is physically separate from your primary data and, preferably, can't be destroyed with mere admin access to the system. The number of sites that have had catastrophic data loss due to relying on mirroring instead of true backups is quite significant.