The 3-2-1 Backup Rule – Why Your Data Will Always Survive (2019)
vmwareblog.org
vmwareblog.org
A slight twist of this: I now have data old enough that accessing it with modern computers is starting to become a challenge. Thankfully I migrated all my, once enormous collection of 50,000 MB on tape, to just be files on my file server, but I'm worried about the longevity of optical media, and now I nervously glance at my collection of even older media ....
You'd lose that bet. 2 TB on LTO Ultrium tape costs under $10 (sometimes under $5 depending on volume of tapes ordered).
Seems like a double standard wherein people are going to pretend housing multiple HDDs costs $0 (unrealistic) but won't even evaluate spending money on a tape drive.
You what doesn't get cryptolockered? Yesterday's tape sitting on the shelf.
If you can move a tape to a shelf, you can unplug a USB cable.
And to your question about having a separate machine, a Pi will be fine.
LTO's combination of cheap media and expensive drives is great for people with rooms full of tapes, but it makes it pretty unattractive for everything else.
You can always feed your LTO2 tape into a LTOx drive and get the data back.
Obviously you have to be fairly patient, but for my use case (once a week on my third tier of backup) it's good and gives peace of mind that I can put in a fireproof safe.
Primary backups are still on the cloud though - S3 Glacier is a good option.
- For home use, it is likely a once-or-twice-in-a-lifetime purchase (try saying that about any other kind of media)
- You don't have to buy a brand new drive of the latest generation at sticker price; used last gen(s) gear from the enterprise works just as well.
- Perhaps most importantly, what's your data worth? I don't know about you guys, but I've got photos and documents that are irreplaceable.
Tapes beat all other storage media on $/GB, reliability over time, and arguably durability, all of which are the most important factors for offline backup.
$500 comes out in the wash over a decade or two. That's the kind of time scales we're talking about. Yes, it's not cheap and easy consumer electronics you can buy off the shelf at Walmart, but it's not unreasonable either.
Two important things to note:
1. LTO5 and onwards have a feature called LTFS, which allows you to address the tape as an ordinary file system. Before that, you were limited to purpose-built tape tools.
2. Generally, tapes can be read in drives 2 generations newer, and written in drives 1 generation newer. This rule was broken with LTO8, due to new tape composition. Starting here, it's 1 generation newer read/write.
Debian + old HPE server + tape drive + tar. Job done.
If the devices are not too large or ridiculously expensive, than it definitely makes sense to use tape as backup. Hard disks are convenient but fail at inopportune times.
I once has to restore a DB from backups just to discover that the backups from 2 days ago were corrupted. Fortunately backup t-3 days worked. And I had to do some binlog mongering to recover the rest of the data.
Test your backups people. They WILL fail.
https://www.joelonsoftware.com/2009/12/14/lets-stop-talking-...
He was saying similar things at that time (e.g., “Backups are worthless, Restores are priceless”), and I’m certain he wasn’t the first.
But you’d have to ask him who was saying things like that before he did.
Disclaimer: Curtis was a co-worker of mine at the time, and I was a technical reviewer of the book in question.
If it was easy then people wouldn't have to be reminded to do it—it'd just get done. :)
One way could be to pre-compute checksums of the files in question and then verify them periodically:
* https://packages.debian.org/search?keywords=aide
It's also why people suggest using ZFS: it has checksums on the live data to make sure bits aren't flipped, but those checksums can then be replicated to remote copies (zfs send/recv) to make sure your copies also coherent. A service like rsync.net allows remote ZFS data sets—that can also be encrypted.
Unless you've tested and proven to your self you can bring up a working system from your backups, you've only done the first half of your disaster recover work.
For me, that means restoring onto your cold spare hardware (either identical or similar-enough for you), shutting down your system pulling it's drives and replacing them with blank ones and restoring onto that (but be careful you can be sure you don't need that specific hardware/firmware/peripherals - cause in 18 months time when someone backs a truck into your office and load all your electronics into it you might not be able to get an identical system), or restoring a running system onto newly provisioned instances on cloud-provider-of-choice.
Prove to yourself at least, that you can have business continuity in a know amount of time once you pull the pin on your DR/restore-from-backups plan.
Of course, it will depend on several parameters, but in my experience, doing a thorough test for a random subset of items is often more economical than a half-assed test of all items.
As a bonus, the random sampling will let you infer things about the totality of all items. (As opposed to any other scheme for selecting which items to test.) So once you've run 27 tests and only one failed, you know at least 85 % of your backups work. At 1/20th of the cost of testing them all, this is a good deal on information.
That could be prohibitively expensive. For personal use, each time I buy a new computer (desktop / notebook alternating) I use it as an opportunity to test my restore procedures. I restore my most recent backup to the new computer, then verify that it operates the same, and has all the same data stores, as my next newest computer (the one whose backup I restored).
I somehow thought the company making the M-DISCs was dead but apparently Verbatim (and others) still sells BluRay discs labelled as "M-DISC" and they're compatible with many BluRay readers/writers.
Typically if the surface of the disc is gold and darkens with writing, it's an LTH disc.
I have that problem!
I have files in my Documents folder going back to the 90s and Classic macOS!
Same. And while I can be sure I have them safely stored, I know for sure I have files I can no longer use because I don't have a way of running the applications needed to work with them. I must have _dozens_ gigabytes worth of copies of Zip disks and Syquest cartridges full of Macromedia Flash projects and Mac OS9 Filemaker Pro databases.
That's probably overrated and you shouldn't feel bad about not implementing it. The justification for it is:
>you may lose them due to the same hardware issues
buying different hard drive models/manufacturers/batches achieves the same thing without the risk of having data stored on dinosaur media.
I wouldn't trust multi terabyte drives in raid 5 sets to not suffer cascading failures during a rebuild after a single drive failure/replacement.
I'd consider two file servers or NAS boxes with the same type/brand/model of drives in raid 5, to be "the same media type" in the context of that advice.
>> The thing is, while keeping data on the same storage media, you may lose them due to the same hardware issues. In other words, you may lose two copies in the same accident. That’s why you should always combine media.
For "same media" the key part of this is storing the data on two DIFFERENT DEVICES of different types, they both could be hard drives, but they need to be in isolated systems of different types (say a Windows System and a Synology NAS, or a FreeNAS Storage Appliance and a Linux Server, etc)
The rule came about because people would use the same tape library, or have multiple copies on the same SAN (often times they would have a SAN Cluster of 2 or more devices that act as 1, and because they had 2 "devices" they felt they were protected.
Disparate media type (HDD, Tape, DVD, etc) may have some advantaged but as long as you are putting 2 copies on say your Home Desktop, and a NAS you have satisfied 2 media's even if both are using Hard drives
That said today the most common way for a person fill offsite and separate media is to use a Cloud Backup of some kind.
Which has a different failure mode to disks/NAS/fileservers.
AWS/Dropbox/Google/BackBlaze et al. are all at risk from credit card and/or financial failures. I can go broke and be unable to pay my bills, but my hard drives will still store my data.
Google and AWS (to a lesser extent) are also occasionally subject to arbitrary and capricious account revocation. Google specifically worries me there, there are way too many stories of people have problems with, say, they Google Play developer account or their YouTube account, and finding they're cut off from their gmail and google drive access, with pretty much zero way if fixing it unless you have a million twitter followers or a Google insider to go to bat for you.
3 Copies, 2 Media, 1 off site.
One should always have a local copy of their backups.
Lesson learned! I now still had several intact tape backups of my data, but no functional device that could read and restore from them.
> can be considered statistically independent too since it is connected over the network and may survive if something bad happens to a part of your infrastructure.
So my interpretation is: e.g. a HDD connected over SATA, an external HDD over USB and a HDD on NAS are all considered different media types even though ultimately they are all based on spinning hard drives.
For my home stuff, I'm comfortable enough with having a backup on (raid 1 mirrored spinning rust) on my media server, plus a copy of that backup (also on spinning rust) on an external USB hard drive. The chances of simultaneous corruption of those - even though they're both spinning rust - seems "low enough" for me.
(I also have another raid 1 pair of spinning rust that powers up once a week very early Monday morning and rsyncs the media server backup directory, then powers back down when it's done, to mitigate against fat-finger-fuckups and/or home network p0wnage. Somebody who p0wns something inside my network might find the cronjob or shell scripts that run the wifi powerpoint those drives use on and off, but at least I'm down to an hour or so a week where that copy might get cryptolockered...)
There's LTO (Linear Tape Open), LTO-8 has 12 TB of raw capacity and costs about 65 - 70 €. The newest variant (not yet _that_ widely available), LTO-9 has 18 TB.
You need a few to handle rotation, but for most that means at max 10 plus one new per year for the long term archive.
We provide Tape support in our open source Proxmox Backup Server product, it can handle single drives and drive robots (with auto exchangers that lessen the work on tape rotation), the data is deduplicated, compressed and optionally encrypted.
https://pbs.proxmox.com/docs/tape-backup.html
PBS can also efficiently mirror to remotes:
https://pbs.proxmox.com/docs/managing-remotes.html
Check the introduction/main feature section for more info if you're interested: https://pbs.proxmox.com/docs/introduction.html
But yes, forking over a 2 to 4k for a new LTO-7+ changer isn't a small investment for private use, for a company it's IMO a no-brainer though.
FWIW: You can get them often cheaper, e.g., at ebay or whatever your local online reseller space is. For example, one can get LTO-5 ones here for 300 - 400 bucks and LTO-6 for 600 to 800. While LTO-6 only has 2.5 TB of space you can also create tape-sets (e.g., spanning multiple tapes) and for the more important data it may even be enough.
No, that's not really a concern currently, and I actually never heard that in combination with LTO - which is relatively modern.
I use a crazy long passphrase to encrypt my backups, but should I forget - it is also printed on archival paper inside a sealed envelope in a friends safe deposit box (I also have a copy of his backup passphrase for mutually assured destruction :)).
Also, every once in a while run a fire drill and actually restore something from each of your backups. This is when you find out the rsync job has been stuck for the last 80 days.
I just wish that storing things like keys on paper was easier.
At CoreOS we put some keys on printed QR codes and scanned them with an airgapped laptop every 90 days to confirm the keys were safe.
I agree. But I've never been allowed to run a fire drill. Rebuilding a network from bare tin is obviously expensive, but not as expensive as losing the business.
And then there's the sheer stress of being responsible for the backups, but not being able to test bare-metal recovery.
I'm interested in backup ("what kind of weirdo is this!"), but that wasn't a fun responsibility.
Sometime your job is not to be able to take the backups, but to take the blame when a restore can't be done. If you think that might be you, time to polish up that resume...
Especially for cloud-based businesses, refusing to let you spend ~8 hours of production platform costs to run a full platform rebuild fire drill is insane, that's about 0.1% of your annual prod AWS budget, and 0.5% of you (or your team's) annual time.
The boss owned the company; I think it was shortsighted of him. Being able to blame me wouldn't have done him much good, once his firm had gone down the tubes, and I'd moved on. But it was his business to lose, and he knew what his margins were - I didn't.
He didn't like me much, mainly because he was a control freak, and I knew his systems much better than he did - I had the control that he craved. Control freaks shouldn't hire people that know more than they do.
I have been involved or aware of a few nightmares where an expected loss resulted in all kinds of pain and problems for the family has they lost access to critical accounts, financial data, and other things that required lengthy interactions with government agencies, banks, etc that would not have been needed if the data would have been able to be recovered from the systems
There are very clear directions to my friend for when it is acceptable to access my stuff, and very clear consequences for malice.
A quick backup verification is not enough - I learnt this the hard way and lost 9 months of data only after I did full annual DR test. The backups were setup and configured by the (very well known) manufacturer of the backup software. They screwed up on the DNS name
I have two friends who've given me private key fragments (from a Shamirs Secret Sharing setup) with the explanation "If anything happens to me, you'll work out who you need to talk to and what you'll need to do."
I haven't done that myself, because I don't have any need/desire for anybody to decrypt my backups once I'm not around. That might change if I end up with dependants one day.
Fire drills are important, but I get a Telegram notification every day when my rsync jobs complete. The notification is muted, but I see it in the list every morning when I use Telegram to text my girlfriend. If it wasn't there, I'd immediately know something wasn't working.
Why not just remember the password and perform a regular fire drill decryption to ensure you won't ever forget it?
“Do not backup everything!”
- repeatedly clean your data
- and discard redundancy and garbage before adding to backup set
- otherwise you’ll just to be creating a backup set that’ll turn into a pile of “mostly” garbage soon
- and it’ll just be GBs and TBs of “too much” - essentially useless - and costing more as well.
Keep your storage footprint in check.
Here's the read out from my last Borg backup which says that I'm storing 2.5Tb worth of data in 59Gb
------------------------------------------------------------------------------
Original size Compressed size Deduplicated size
This archive: 220.28 GB 66.75 GB 591.11 MB
All archives: 2.55 TB 777.85 GB 59.44 GB
Unique chunks Total chunks
Chunk index: 234389 4706300
------------------------------------------------------------------------------I.e. I backup my photo collection religiously. I keep all my photos in the cloud, and have a machine synchronizing photos locally. That machine then makes 2 backups, one local, and one to another cloud. The same applies to documents, mostly because they're highly compressible and don't take up much space.
When it comes to media backups, i honestely don't care if my iTunes library got wiped out. Most of it has been purchased, so can (hopefully) be downloaded again, and the rest has been ripped from CDs that still reside somewhere in my attic. So inconvenient to lose, but not exactly critical.
I also don't make "full computer backups". If something breaks i can just as easily reinstall the computer/applications and restore my documents/photos.
As for photos, i also burn identical M-disc BDXL media every so often, containing the photos taken since last archive date, and store the copies in different locations. They're low cost, low maintenance, and while not "spinning rust" cheap, they're still within $12/100GB, and unlike spinning rust and cloud, it's a one time cost.
What if the source for the backup was already corrupt or broken in some way? If you only have one backup, then your backup is corrupt too.
I was taught grandfather-father-son back in the 80s; still three levels of backup, but they're different generations. That fitted the kinds of media available then, but it doesn't really map to modern equipment. I've struggled to work out a backup scheme that is equally adaptable to the needs of a small business, a home network or an individual.
Ironically, it's hardest for the individual; a modern business is finished if it loses all its data. For an individual (or even a hobby network), total data-loss is painful, but not usually an existential risk. So it's harder to justify keeping everything in triplicate.
I would like to have a grandfather, a week old; a father, a day old; and a son, being last night's backup. All on different media, amd ideally not connected to the source machine.
Ransomware is the #1 thing an enterprise will use a Backup to recover from, the second most common is accidentally deletion.
Binary corruption is very very very very very rare and with all the systems in place to prevent it, is not even something I even think about anymore. I worry about Ransomware, and users doing stupid things
So, how do you detect if you have ransomwared data? My take is that a backup can't help with that. It can help with restoring obviously.
I agree with you problem description and approach. I'm less sure I agree with your terminology choice.
I use a software raid 1 pair of external usb drives on a wifi powerpoint to take a once a week snapshot of my backups, and the drives are only powered up for ~1 hour a week. Not guaranteed protection agains ransomware, or me fat fingering a "sudo rm -rf / tmp", but "good enough" for my home/personal stuff. I also have monthly and yearly snapshots of my backups in AWS Glacier. I sometimes colloquially refer to all that as "backups", but I can see your parent posters point that I've gone way beyond what "backups" covers in this paragraph.
DR is a set of policies and procedures that cover backups. That are inextricably linked
You use backups to perform a diaster recovery.
I personally dont believe you can talk about backups with out including DR nor can you talk about DR with out talking about backups
They are linked, but you can do backups without doing disaster recovery (for example, just turning on TimeMachine backups to a usb drive on your Mac is “backups” without being “disaster recovery”), but you can’t do disaster recovery without backups. But you also need archives, and retention policies, and recovery procedures and test plans, and training and practice for the people responsible for DR, and hardware/site/network disaster recovery plans and resources, and recovery time objectives and recovery point objectives, and a whole bunch of other “not directly backup” related stuff.
But I admit that colloquially that might all be assumed in certain contexts to be “backups”, but that’s a probably dangerous assumption unless everybody in that discussion is totally on the same page about just how much of that related disaster recovery stuff is actually in place and ready.
A proper backup tool should help you keep several versions of your data without using a proportional amount of space, by using some form of deduplication. I use borg backup for my backup, and I can go back to any day in the past three years and get an old copy of any document (as long as I saved it on my disk for more than one day, since I do daily backups)
You can also setup "append only" backups, if you are worried that somebody may willingly try to destroy your old copies
I think about backups in terms of blast radius. 1) The local machine has the working copy of data and a local backup as permitted by free space. The smallest blast radius where I lose data is "my laptop hard drive fails". 2) My external drive has another backup. The new blast radius is "my house burns down". 3) I maintain a cloud backup. The new blast radius is "a catastrophe on a global scale".
Any two of these backups can fail and the data is still salvageable.
It exists and is important because many backup strategies are broken and people don't realize it.
For example your own strategy treats a PC's local storage and an external drive as distinct backups when in reality you've only evaluated hardware failures when formulating it and not malicious actors. In 3-2-1, in particular the media type + off-site thing, tries to "trick" you into having a backup which isn't accessible from the same system that it is backing up (i.e. offline backups).
Your backup strategy has been used almost verbatim by multiple institutions who got cryto-locked. The external drive was hit and then the cloud backup service happily synced the now encrypted files. They went from "backed up in three places" to backed up in zero places, and are now calling the cloud provider hoping that their backups had unencrypted copies in them.
3-2-1 isn't simple, but it is good, and that's what it tries to be.
Is that good enough in your point of view?
Thanks
[0] https://www.rsync.net/resources/howto/snapshots.html (not affiliated, just a happy customer)
[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/Versio... (ditto)
For targeted - it often doesn't because often your AWS keys are on the system doing the back up and have permissions to delete items etc.
Of course, this is why S3 allows you to set and object lock rule (ie, 30 days is plenty) that means even you (or your computer) can't go and delete those online backups.
That's because a targeted attack with a ransomware could gain access to your servers and wait 30 days while silently encrypting your backups until the 30th day, when the attackers could just complete the attack encrypting the rest of the files and showing the message.
So my recommendation for extremely critical data would be: 1) Test whole data thoroughly at least once during the object lock period. 2) Setup an automatic task that retrieves X random data every day (or the longest period of time you can afford to lose it) and perform checks with checksums or other methods. If something is corrupt and/or encrypted you will realize before it is too late.
Just a reminder - snapshots at rsync.net are immutable.[1]
Even if Mallory gains access to all of your access credentials, she cannot delete/change the snapshots in your rsync.net account.
For cloud, I use a time4vps storage server. I rsync some data and use borg for more sensitive stuff.
Of course if the provider is locked then it's a moot point.
Presumably the rule was invented before ransomware was a thing, so perhaps it gets a pass for not anticipating versions, but yeah: the rule of thumb for the modern world probably includes something about backup versions.
Veeam for example (other backup systems are available) have several "immutable offerings" whereby you get backups that can't be deleted for a while. These are not appliances but configuration recipes. One of them is a generic Linux box with XFS with reflinks, a particular set of file perms and what they call a one off credential (a sort of app password I think), forward incrementals and a few other things. Get it all in line then even if a baddy gets in they can't delete your backups for a while. This does require detection inside say a week. The reflinks thing is really useful for storage space and things like "synthetic fulls" where you roll up incrementals into the last full backup.
You can also use AWS etc object storage for immutable. In all of these things you can lengthen the timescales available by applying more cash.
For me: tape in a safe is the only decent way to protect against ransomware. Even then some tape hygiene is needed!
I think there's an element that is as important as keeping a backup safe against a state-wide wildfire and that's AUTOMATION. If it isn't automated then chances are that your backup is very old once you finally need it.
far more simpler solutions exist that don't provide anything like reliability.
Things area more nuanced these days than when this "rule" was first formulated, but I suspect it's still true that the vast majority of people would be far better off with following it than whatever they are doing now. Doubly true of personal use.
I only use cloud backups (Glacier and Google Workspace), I gave up on off-site drives as they end up too far away to be easily/consistently updated, or close enough they are in the same disaster area I am in (Earthquake zone).
Houston flooding is another good example.
Personally, I downsized. I made peace with myself that I can live without the gobsmack amount of data I have if it were to be lost tomorrow. I pared down to a small set of data (less than 1GB) that is critical to me. That data is synced to various devices with syncthing (includes even my cellphone!), and then I use restic to two different cloud storage providers. When I'm bored I do an independent, standalone export from cloud storage.
Since I "restore" this backup pretty frequently just for day to day living (ie, doing taxes) I'm also pretty sure it's accessible.
I do pay for versioning for the online sync of this, and I do a period S3 object lock copy (30 days). For me, that's good enough.
Archive links:
https://web.archive.org/web/20211001064106/https://www.vmwar...
3 backups
2 different media types
1 off-site
Having several drives is not enough. I used to keep my important data replicated on 3 drives, from different brands, different capacities, 1 internal, 2 external.
One day the internal drive failed, the next day one of the external drives also failed. So... I panicked, shut everything down and bought a brand new hard drive (1tb). While copying from the third drive it also stopped working. So, I had a triple drive failure. I managed to recover most of my data by freezing the external drives and copying from them (until they heated up and had to freeze them again).
Pretty high for any parity-based multi-disk system. The remaining disks get stress tested when a disk fails and you need to copy everything over to the replacement. It’s why RAID5 is no longer sufficient with today’s disk sizes and why a RAID10 (which can only lose any one disk plus a specific other disk) is actually real-world safer than RAID6 (which can lose _any_ two disks).
There is also RAIDZ3 for those who want more safety, which is ZFS triple parity.
- One live copy of data on my laptop
- One copy on external hard disk, updated ~daily, on-site (home)
- One copy on external hard disk, different brand and age, updated ~bi-weekly, off-site (work)
- Immutable copies on optical media, persisted ~once a month, on-site
Data on the laptop was lost due to operational error: I fat-fingered a command and destroyed the partition table and part of the leading data on disk. Being a reasonably fast disk this ate a lot of structurally critical data quickly. Recovering the filesystem would be really hard, but I had a two days old backup, so didn't think much of it.
Now, to the local backup. I booted up a live cd, rebuilt a partition layout, plugged the disk in, and started restoring data. Reboot, and it seemed to work, mostly, but some things that should were not. Immediately jumped to look at the recovered data, it was severely corrupted. Diffed some files and compared to the "originals" (i.e from the backup) and they were identical: data on the backup disk was hopelessly mangled even though the hardware was fine. A cursory analysis seemed to highlight a software bug (filesystem code? drive firmware? whatever, the issue had some logical consistency to it that made it obvious it was unrecoverable, which was my sole goal at this point). And I just restored it over what remained of perfectly valid - if difficultly reachable - data, essentially scrubbing the laptop disk. Sweat was starting to build up.
Okay, optical copies were next, even if older. Surely this would get my heart rate down. I put the disk in, closed the tray, and heard the sound of a rattling helicopter. I stored the disk in a closet which I thought would be safe, but it turns out the hot water pipes for the flat above were running behind a thin wall, which build up enough heat over time inside the closet to slowly warp the disk. Well, one of them, because I was paranoid enough to have three disks for a 3 month rotation; but while the other disks were geometrically fine (maybe due to being a different brand), they were stored for a longer amount of time and their data suffered much bitrot. This was going to be a long Sunday.
Back at work on Monday, the final disk immediately emitted an ominous clicking noise right away. Shortly after it snapped, never to power on again. I could maybe recover data straight from the platters if I sent it to some firm for a hefty pile of cash, which I had none at hand, neither at that time nor in the foreseeable future.
So, in order I experienced: an operational error, a logical error, an environmental issue, and a hardware fault. Luck had it that I had a second computer temporarily lent to me, which I toyed with and where some of my most recent work files turned out to lie from a week before, so I could resume putting food on the table quickly. No amount of hackery was able to restore any meaningful data, so I lost about 10 years of digital photos and older work archives.
Psychologically it was fairly interesting, because I thought I would be enraged at myself for multiple reasons, but the perspective of such ridiculous odds of this happening turned the whole thing into a very contemplative experience shortly after.
I've never had a total catastrophe, but I have had a chain of independently-unlikely faults that combined together to create a once-in-a-lifetime disaster.
This has happened to me many times during my life.
From experience, hard drives seem safe on the order of years; I've spun drives back up from the early 2000s and they are fully intact. The lifespan of burned optical media is/was counted in minutes. Various flash memory is somewhere in between - I've had all kinda of cheap flash drives die.
Even with the multiple media forms they need to be refreshed at some frequency, and I don't know how often that should be.
Trying to find an authorative source (loc.gov, archive.org, etc.), I found this, which is not a full answer, but gets into interesting details: "Table 2 - the relative stability of optical disc formats" [0] -- from >100 years down to ... 5-10!
[0] https://www.canada.ca/en/conservation-institute/services/con...
> There are Blu-ray Discs specifically engineered for long-term archiving and have BER guarantees (anomalous bit read per gb of data stored per year archived or something).
Now these are obviously simulated numbers (the tech isn’t even old enough to test) but it’s a start: it means people are at least considering the right questions.
I wouldn’t write sensitive to a Blu-ray directly unless it was the kind of data where a bit-flip is not a huge deal (eg a backup of users’ profile images where there are many small files, a bit flip affects the content but doesn’t compromise the overall data, errors aren’t cascading, etc). There’s already bit error correction baked into the analog <-> digital transition layer but it’s not great - but fortunately efficient bit error correction at the file level has been a thing since before binaries on Usenet. Stick some PAR2 files on the Blu-ray or even serve your backup as a Blu-ray Disc plus a DVD-R stuffed to the brim with PAR2 data for the former (I prefer the first approach).
It's crazy how the seemingly easiest most basic security/backup advice is so easy to give, and so hard to actually do. 3-2-1, so easy to teach and remember! In reality, at any kind of scale, not so easy to do.
I am constantly reminded just how hard every aspect of security really is to do. Even for the little/basic stuff.
1. Local disk
2. Time machine ssd
3. Backblaze backup agent
For cloud stuff the services make it so hard. I have been working on a service off and on to backup Google photos to an SD card and then mail it to folks. And the amount of limitations, rate limiting, and random errors out of Google's API can be frustrating.
Not through Google's APIs. But you can, through adversarial interoperability.
For me, the only sane thing to do is partitioning.
My first group would be data that if I were to lose it would cause great pain. I keep the size of this group as small as possible. A few hundred megs or less. You have a live copy, a backup on a USB thumb stick or drive, and a copy you email someone or snail mail a USB drive. It's simple to deal with. It has to be, because it's critical.
My second group is data that is important but not a serious threat to me. Photos and videos, mostly. This second group is where the headache starts and logistics, cost, and time become an issue. Offsite backup is either running to the bank deposit box (time consuming), or upload to cloud (also time consuming, and expensive). Containing the bit rot becomes a futile exercise. Especially considering most people aren't running ECC RAM and end-to-end ZFS with redundancy (for recovery) requires significant expertise and time. Parchive files are the best bet for most people.
Finally, my third group. I have a NAS with a simple mirrored ZFS setup. Two huge drives. I'll probably add a 3rd drive for added redundancy. There is no backup. This data I don't care that much about. I'd hate to lose it. But I'd hate backing it up much more. I don't live for my data, my data lives for me.
You have to ruthlessly prune data that you care to keep. Just like the burden of owning a boat or an overly large house, there is a burden to too much data you care about. The mistake a lot of people make is treating all data the same. Then they end up with terabytes of data of unequal importance and get sloppy protecting the tiny amount of data that truly matters.
My first service order way back when was to a engineering office who were in a panic because their primary CAD file server had failed right in the middle of a deadline submittal.
I got there when they had just pulled out an identical file server from some other department. They had 'everything' backed up nightly on to a Zip drive. Nice, I thought. So, I plug in the Zip drive and can see all these image files created by their backup software. I ask for the backup app install files and they say its kept in a directory -- on the failed server !!!
I should mention this was not in the USA and resulted in a 6+hr long international 42kbps zmodem download session from some random BBS server of the best guess software product and version.
I still have one of the failed HDDs from that server as a paper weight on my desk.
PS. We got it working in time and the obvious moral of the story is to always test your backup systems (and that goes for both failover and fallback)
I particularly like to do this when showing someone around. You whitter on about power and then you swiftly turn the key on the distribution panel and it makes a satisfying clunk. The status lights switch around and the UPSs start beeping. Then they stop. There is a barely perceptible hum from the genny in the boiler room, three walls (one of them a firewall in the traditional sense) away. It is extremely satisfying to do and customers appreciate it - they have all jumped a bit when they first see it. The UPSs obviously (monitored and fortnightly tested) have enough runtime to switch back if something fails on the genny.
At a second location there is just enough IT gear to run my company for a while. Replication via Veeam - 'nuff said.
Every couple of weeks, Veeam fires up a small part of my company's IT from backups/replicas at the second location. It runs some scripts which tests some functionality and sends out some status reports via email.
I still worry about business continuity. There is way more to it than just backups. However, get backups sorted out first. Nine months back a customer had a fire on the production floor. They are still running out of our place, 50 miles away. We set up offsite backups to our place only the week before: "There but for the Grace of God, go I" etc etc. I had them up and running inside three hours from a standing start. Rather lucky there really - they have a functioning business and I have a success story rather than a potential lawsuit from an ex-customer's liquidators trying to claw back some loot.
I'm not affiliated with them, just a happy customer enjoying massive cost savings.
> 80% Less than Amazon S3
Why would I use a service that is 80% less reliable than Amazon S3? And if they mean something other than reliability - such as uptime, price, or transfer speed - then they should state it clearly.That line could easily be a decoy to ward off problem customers who don’t know how to use APIs, read specs, etc and would be high customer service load.
In any case, thanks for mentioning the service. I'm quite happy with [Boto](https://github.com/boto/boto3) for interfacing with Glacier but it is good to know about alternatives.
Basically anything that gets uploaded is billed as if it is was stored for 90 days minimum, even if deleted, overwritten etc.
An additional difference is that the egress is limited to the size of the buckets. So if you store 1 TB your egress should not be more than 1TB in a month.
You need to add the time dimension and have a monthly snapshot for data that is older than a year, weekly snapshot for data within last 1 year, and daily snapshot for the last 90 days.
If backups are not tested by actually retrieving data on a regular basis, you might get a nasty surprise.
The Legato NetWorker bug which resulted in 64 bit XFS inodes not getting backed up bit Pixar very hard back in the day (LGTpa40680).
- integrated tape backup (LTO-4 or newer), this allows to fulfill the two different medium rule and may also help on the offsite one (e.g., but a tape in a safe at CEO's home once a week or so)
- efficiently mirroring of backup data to remote PBS (offsite copy)
- client-side encryption, because well, backups are good but maybe not so if they're leaked
- file-level and block level backup - disclaimer the former currently only for Linux clients and the latter can handle anything but also runs on Linux (VMs are used)
The introduction/main feature section of the docs contain more info, if you're interested: https://pbs.proxmox.com/docs/introduction.html
If you have your non-Linux workload contained in VMs and maybe even already use Proxmox VE for that it's really covering safe and painless self-hosted backup needs.
A script touches the copy mail store (I currently use Zimbra for a mail server) to check it is up, and that the last item “received” is no older than 24 hours, and emails me to say all is well. If I don't get that daily mail, the restore has failed somehow. I keep meaning to add an extra layer that checks for the daily mail and sends me an SMS if it doesn't arrive, just in case I don't notice. The copy VMs don't need to be high spec, though in theory if the main machines die I could just switch other to them in minutes by DNS+firewall updates (and by giving them a bit more CPU+RAM allocation), just enough that the services start and run fast enough for the occasional manual paranoia check, and they aren't accessible to the outside world (but could be, if I needed to make that switch in a DR situation).
Similar for the little web servers I still run and key parts of general data backup (the parts that can't be reobtained at all if both originals and backups are lost): restore from backup after the update of the backups is due to finish, checksum everything older a few hours on both sides, and send a mail if the checksums match. That sometimes give false negatives, if a file is updated between backup and checksum, I have a couple of possible fixes for that in mind but it only happens if I'm working silly late due to when the checks run so I've not been bothered enough to implement one.
(Side note, I also see the same comments about the shoes I want from the company I've bought shoes from forever).
Is this comment overload from naive users in the hard drive sector, or has quality control dropped across the board?
The average rating is probable a better indicator here.
I usually judge based on the manufacturer warranty. If they do 3 years, it's probably pretty good. 5 years means top shelf. Stay away from anything with shorter warrenties.
If there's a backup process you should also test it too. Scenarios like backup job workers are stuck and the system is not actually backing up is not uncommon.
This would take care of 3 and 1, sadly not 2 but 2 is pretty hard.
They offer an amazing number of 9s for durability, it is very unlikely that a single fire will cause harm to your data.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/replic...
https://aws.amazon.com/blogs/aws/new-cross-region-replicatio...
In short: AWS will go down only when you won't care about it anyway
There is also multi region buckets which go across regions.
I use RClone and Duplicacy to backup to S3 Glacier and Google Workspace. Kopia is also another up and comer.
For a simple copy type setup RClone is a good one to look at, by default it copies the files and doesn't do any snapshots, packing, splitting, etc.
They really don't seem to give it any thought at all, even when I explain to them why they need to do this.
Conversely, when they're having network connection issues they don't hesitate to call me, and sometimes in a panic.
I'm going to keep pushing them to get on it on their end though.
I’ve found some content specific solutions that will backup photos or contacts for example, but I’m looking for a fairly streamlined comprehensive solution that would back up everything-messages, contacts, photos, emails, bookmarks, basically iCloud 2 but without the Apple lock in.
... After a frantic search that entailed calling hundreds of IT admins in data centers around the world, Maersk’s desperate administrators finally found one lone surviving domain controller in a remote office—in Ghana. At some point before NotPetya struck, a blackout had knocked the Ghanaian machine offline, and the computer remained disconnected from the network. It thus contained the singular known copy of the company’s domain controller data left untouched by the malware—all thanks to a power outage. “There were a lot of joyous whoops in the office when we found it,” a Maersk administrator says.
When the tense engineers in Maidenhead set up a connection to the Ghana office, however, they found its bandwidth was so thin that it would take days to transmit the several-hundred-gigabyte domain controller backup to the UK. Their next idea: put a Ghanaian staffer on the next plane to London. But none of the West African office’s employees had a British visa.
So the Maidenhead operation arranged for a kind of relay race: One staffer from the Ghana office flew to Nigeria to meet another Maersk employee in the airport to hand off the very precious hard drive. That staffer then boarded the six-and-a-half-hour flight to Heathrow, carrying the keystone of Maersk’s entire recovery process. ...
https://tech.industry-best-practice.com/2018/10/14/the-untol...
Write-once seems valuable in the age of ransomware.
The discs themselves are quite expensive in terms of cost/TB - somewhere around 50-60 GBP from what I've found.
I have more confidence in their ability to not lose data if left on a shelf and forgotten about for several years compared to hard drives. I have the same data stored on a ZFS pool too.
Had 2 backups (1 SD card, 1 HDD) for my "Document" folder.
I was trying to replace my GRUB MBR bootloader for REFind's EFI so I could dualboot on a new laptop (and swap-in my old one's SSD without having to reinsall the whole system: Arch). Unfortunately, the boot partition was too small, and needed to be re-sized from 512MB to 1GB. Foolishly, and in a rush, I thought using gparted to change the partition boundaries (shrink the root partition by 512MB from the beginning, and stretch the boot partition to 1GB from the end) was the answer.
I completely forgot that EXT4 has a superblock at the beginning, so now it was gone, and the root partition was completely unmountable -- and fsck was of no use.
So I scramble to find my backups (to decide whether or not I should figure out how to fix this), and realize that tiny little SD card was missing, and my HDD backup was completely unmountable.
Truly, a major fuck-up.
Thankfully, I didn't write anything to the boot partition, so throwing a Hail Mary and simply resizing the partitions back to their exact original sizes (thankfully x2 my TTS history was useful), allowed the root drive to mount without a hitch.
I was close to losing all of my KeepPassXC passwords and private keys due to shear idiocy.
In the end, I set up "cloud" backups (second storage media type, and long and far away), switched to Debian, and continued on my merry way.
I've made similar mistakes (well, the same effect, but for different reasons. I thought I was over a remote session when I was local and made an entire new partition layout).
`testdisk` was able to introspect what partition boundaries _were_ and rebuild the partition table.
If you're ever in a similar situation and either don't remember the exact boundaries or don't trust yourself to recreate them, check it out!
> switched to Debian
Not to trivialize the wonders of Arch, but the switch to Debian was a major step in improving the Linux-Wife balance. A system that is reliable, and doesn't have many history lines that begin with `sudo vim /etc...` makes for a happy family life.