I lost my data trying to back it up
hamy.io
hamy.io
And how did that work out for you?
If you did believe in online backup solutions, in addition to whatever else, you'd have your data now.
(I understand worrying about quick recovery, booting, etc. But, as an additional method besides hard local copies, it makes sense, as having an extra off site backup of the data is important. Better to have to spend 3 days to manually download and restore a bootable image, than to not have the data at all. And online backup services also allow for automated backups on your backup -- backup-ception).
3 copies, on 2 different media, and 1 offsite.
I also feel like adding, replication is less likely to work correctly the more abstract it tries to be. Compare hardware RAID (pretends to the rest of the computer that there is one disk instead of two) to application-level replication (don't return success until the document has been written successfully to 3 nodes). That first one never works correctly, because there's so much that can go wrong and so little you can do to fix it. The on-disk format depends on some random firmware that is treated as disposable by the manufacturer of your motherboard; with no documentation or source code to be had. The second one pretty much always works because it's simple, and when it breaks, you can just examine the content of each node and figure out what went wrong, because you designed it and it's as simple as it can possibly be. (If you want to not store 3 copies of your data, you add your own Reed-Solomon coding to the data you write, and pick the ratio of data size to lost chunks to suit your needs. That's what RAID-5 is, just abstracted over an entire POSIX filesystem and hundreds of thousands of lines of code with no unit tests.)
I would never use RAID again, as I consider the complexity too high for the benefits. Treat filesystems and disks as disposable and transient; don't try to build a filesystem that's durable.
I've found ZFS 'software' RAID to be very good (RAID isn't a back-up, of course). I think data-integrity is another important component - earlier I had multiple versioned backups in different places etc. and still suffered data-loss due to bitrot (which my backups faithfully propagated).
It seems that you should have 3 copies -- 1 being offsite.
But the 2 I have heard the 2 as 2 formats, 2 devices, 2 mediums, 2 storage types, 2 technologies.
I don't think the idea of 2 is reasonable for anyone except the most hardcore. Am I really expected to buy a tape for my local personal backups?
I think this rule could be updated to something 2 different cloud providers. Or 2 different geographic regions. Drop the 3 and the 1.
If all your backups are to a set of tapes or, say 3 hard drives, and you have a rotation to keep one tape offsite, every copy shares too many traits in common. A fair number of Mac people use Time Machine to do versioned backups to an external drive (or drive array) and use separate software to simply mirror their drive to external drives, periodically swapping the onsite and offsite mirrored drives. To me, physically moving drives around is hardcore, I don't trust anything not automated, but it seems to be not that rare. I'm fairly lazy but my machine is a work laptop so I have a backup drive in the office, one at home, and have a cloud backup service so I have 3-2-1 without much thought or effort.
If backups are solely on hard drives, it's probably best to not use the same make/model purchased at the same time, for fear that they'll all fail within the same time frame.
For example you religiously have three copies of data but they all came from the same tape drive that silently went bad months ago and has been writing garbage. Or you recieved a batch of bad tapes etc...
I've also configured a daily mail with a log about how the backup/snapshot/replication went. Even if it went fine. So if I don't get mail, I know something is up and can fix it asap.
I have never slept so well after doing this with ZFS, even backup for Windows SQL Server with iSCSI volumes works great!
More people should do backup with ZFS!
"The parity RAID code has multiple serious data-loss bugs in it. It should not be used for anything other than testing purposes" [1]
In my setup there is no special network tuning either, it is completely basic MTU is 1500, NO VLAN's, both iSCSI and regular network traffic works great without impacting for each other. Performance is excellent with the cheap Mellanox ConnextX-2 ethernet cards you find on Ebay for 20 USD. There are some really excellent, cheap, silent and with 10G SFP+ ports network switches available too for just 100USD
(Hopefully you've tested recovery procedures, because as the saying goes, "your backups are only as good as your last restore")
ZFS snapshots are atomic. Recovering from a snapshot is the same as recovering from a loss of power. All serious databases don't acknowledge commits until either the transaction log or the container is synced to disk and can handle an unexpected loss of power/reboot/system crash without significant data loss.
However, it is absolutely better to put the database into some kind of backup/write-suspend mode when taking the snapshot.
Shut down the DB, or use a CoW filesystem which can do atomic snapshots (zfs, btrfs, bcachefs, apfs, ...)
Even an email saying 'things went well' doesn't tell you that you can restore your files.
IIRC, the ability to remove a vdev has just recently been added as a feature for ZFS-on-Linux on master, but I do not believe it is released yet.
For example once a week I fire up a PostgreSQL instance on my remote backup server by cloning one of the snapshots and then I use pg_dump so I also have a "traditional" backup of the database. I do not spend time on verifying this dump though, as the ZFS snapshot already worked fine :-) But I sometimes use the dump to test new hardware or a new PostgreSQL version.
At work once I set up data backups, the first thing I do is try and replace the backed up server with a restored server. And you should regularly do that if the data is at all important.
I also take that same set of important stuff and rsync it to a cloud server. I have a TB "storage server" from Time4vps (affiliate link: https://billing.time4vps.eu/?affid=1881). I then use borg to backup my more sensitive personal data to my cloud server and 2nd drive.
If you snapshot a backup to provide a history of backups, then the resulting snapshot is also a backup (albeit not a new one, merely part of the original). However, if what you snapshot is not a backup, then neither is the snapshot.
As a good rule of thumb: If you lose data when the machine suddenly bursts into flames in the most literal sense, then that machine cannot even in the slightest be considered a backup. Not even for home use. While the thought of spontaneous combustion seems drastic, total death situations are very real and not rare at all. Lightning, power surge, power supply failure, drive controller failure, simultaneous disk failure, filesystem bugs, and the list goes.
A good backup for things that truly matter is also resistant to the entire building being lost in the flames. For ZFS, the obvious choice would be to `zfs send` to replicate to a machine in a remote location location, but beware of the incremental send hole birth issues.
If you do not use ZFS, or do not have multiple ZFS machines for replication, look at restic (https://restic.net/) and rclone (https://rclone.org/). Restic can through rclone do nice, incremental (snapshot based), deduplicated and encrypted backup to anything from google drive to your toaster. For small payloads (<5TB), consumer cloud solutions such as OneDrive ends up being quite affordable.
> I keep 180 days with snapshots on the remote backup and 30 days for the local.
Hardware RAID (even fake ones like Intel) are time bombs, either you have fun configurations like this or the controller dies and no replacement exists.
Unless you'd absolutely need a HW RAID, I'd go for a Software RAID that mounts and works natively in Linux. I can put the harddrives into a fully foreign computer and it would work.
I’d also argue that for most individuals who don’t really need to worry about an hour or two of downtime, it’s a lot more useful to have good (multiple) current backups than potentially complex RAID configurations.
Making independent devices into a unified device is a concept that is the polar opposite to a backup.
Your example is also why people say that two backups are needed. Once your live system fails, if you only had one backup, you are now left with none. What was a backup is now your only copy, leaving you very vulnerable.
Another identical (hardware-wise) cold spare was already in the rack right below the unlucky one, but it wouldn't accept the transplant license dongle because of service tag mismatch. First I had to install Windows 95 just to run the service tag changer tool and flash the firmware on the spare controller to match the original before it recognized the old disks as an existing array.
In around one week we were back up and running (this was a non-workplace environment so there were no concerned execs bugging us for status updates every fifteen minutes). Had we been using mdraid, we could just have inserted the disks in just about any commodity PC hardware. In hindsight, just restoring the week-old backups would have been easier.
Lesson learned: even if you are using hardware RAID and have an ostensibly identical hardware spare available, it won't necessarily save you.
Raid is useful for availability as well as performance though.
Also while we’re discussing raid, the price/gb of a mirrored set of larger disks vs a raid 5/6/10 of smaller disks (for the same total capacity) is always worth checking - usually comes out in favour of mirrored pair in my experience which also has better performance (especially when resilvering) and resilience and lower energy costs.
Also for data integrity, with something like ZFS.
For home desktops, that is. Server end is different.
That required some great caution.
My worst horror story wasn't really mine. I did desktop support in my first job. The library had their own system and their own support arrangements but I had a good relationship with them and would advise from time to time. One day they said the PC running their archive management system wouldn't boot. Could I take a quick look before they called the company on their support contract? Sure enough the PC wouldn't boot, the hard drive had died. I asked if they had recent backups? Oh yes the backups were fine, in fact it had got much better recently. Before the backups used to take an hour, but for for the last few weeks it only took 2 minutes. They only had 2 weeks of tapes on rotation.
Needless to say, the backups hadn't been running at all and failed silently, probably due to early symptoms of the hard drive failure. They lost everything.
I think the Ubuntu live image is therefore unsuitable for this kind of work. What was needed here is a more specialist "rescue" system image that has a "forensic" mode. For example, Finnix has this: https://www.finnix.org/Forensics
- Your data is at most risk from yourself. Other risks such as hackers, burlgars, fire or hardware failures aren't nearly as dangerous to your data as you are.
- Backup is HARD. Have someone else do it.
- Mirroring is fine for one first copy. Whether that copy is raid or just another NAS or something, but there should always be a proper BACKUP too.
- To be a proper backup it should be 1) Write only/snapshotting, so even if you remove all your local files, the deltions aren't mirrored. 2) Have long retention. When you mess up a document you want to be able to notice 2 years later and just fetch that version 3) Be off site so theft and fire doesn't risk the primary copy and the backups.
If OP had iterated on having a backup system BEFORE adding important data, they wouldn't have consequences to trying to implement it. At the very least, do an online, filesystem level backup of the things that are important before trying to block level disk copy things. Run a database dump, rsync it to an external location, put important things in version control and push to a remote origin. And rsync or tarball your home directory and any other important directories(whole system if possible) and push it out somewhere, then and ONLY then should you feel semi-comfortable to start messing around with RAID settings/LVM/fdisk/ etc.
Mdadm is far more resilient, and I would never go with hardware RAID again without full vendor support.
I do leverage a few things though... for the most part, I keep stuff on my nas. Projects are in git remotes as well as several local copies. Serial numbers, access keys etc are encrypted and stored in dropbox and copied to google drive.
These days since most of my work is in source control, I'm less worried about that aspect. That said, however, RAID is not backup.
A backup is a copy of data on hardware separate from the live system.
A good backup is a backup in in a remote location.
A backup strategy is to maintain a good backup.
A good backup strategy is to have routine restore tests, to make sure that the backups serve their purpose.
A great backup strategy requires at least 2 good backups.
The location thing is trickier, and while you can't have a any backup on the same device, you can have a bad backup on the same location. A bad backup is better than no backup by a long shot.
There's lots of fancy tech words in that article. Before you use ANY piece of backup tech you do not 100.0000000000% understand, back up your data with tech you COMPLETELY and UTTERLY understand how to use.
This likely means a simple copy to an external drive.
After that, set up your first automated backup solution. Test that it works.
After that, set up your second automated backup solution. Test that it works.
For the especially paranoid (me), you can still do manual backups using tech you completely and utterly understand.
Separately, having at least a partial online backup would have minimized the issues caused by the offline backup.
Making a system backup like this should just be about imaging the raw disks, regardless of OS, filesystem, raid arrangement, etc. sda is sda is sda, period.
Keep it simple and you drastically reduce the likelihood of making a mistake like this. A copy of Knoppix and dd will do the job just fine. Paragon Software, Clonezilla, ntfs-3g, mdadm -- distractions.
Even better, if you keep the fundamental units of your system simple (a raw image of sda), you maximize tool compatibility / availability and your chances of recovering when things aren't going to plan.
For complete system backup on Windows use Acronis [0], on MacOS use Carbon Copy Cloner [1] and on Linux Clonezilla [2].
Also never forget the 3-2-1 rule of backups [3], 3 copies, 2 local copies on different mediums and 1 copy offsite (cloud, remote) :-)
[0] https://www.acronis.com/en-us/personal/computer-backup/
[3] https://www.backblaze.com/blog/the-3-2-1-backup-strategy/
Going by the names and the descriptions I've read, the overloaded term 'backup' only applies to Carbon Copy Cloner and Clonezilla in the sense that they make a copy of your storage, so don't protect against corruption (such as ransomware, or you accidentally corrupting files, or deleting files...) Please let me know if I'm wrong on this. Of course you can do real backups with a clone-only tool if you rotate the backup media yourself, but I prefer this aspect to be handled automatically.
For Windows, a similar strategy is possible, but I consider my Windows box disposable (only gaming), so don't have much recent experience.
Unless your NAS is somewhere (not your home's internal wifi network but a WLAN). Consider a Carbonite (or similar) that has infinite backup space to make sure you have the 'offsite' covered.
Honestly though, low effort is very important for me. After all, any solution is simply reducing the odds, not completely eliminating them.
Carbonite backups up 24/7 so I have even the latest changes. The only painful thing would be to find a new/similar laptop to restore my data.
If it would be a different brand laptop, then I would have to go through the pain to setup most stuff. Chances are that restoring to a new laptop may work right off the bat, even if I will have to spend X hours finding appropriate drivers (better than setting up everything from scratch).
It has saved me a few times. In particular it saved my ass when my brand new OCZ SSD died on me on day 2. Due to being full-disk backups, I was back up and running in 30 minutes on a spare HDD with no essential data loss.
I am a bit conservative when it comes to backups and I do not trust myself very much. So I am doing backups on the block device level to avoid not copying certain file types, file attributes or whatever feature the filesystem in question posses.
In order to avoid, risky OS reboots, I am using lvm snapshots like this:
# sync the filesystem to disk
sync
# create a snapshot with a 10GB buffer
lvcreate -L10G -s -n "$snapshot" "/dev/vg/$lv"
# write the backup to a backup disk
dd if="/dev/vg/$snapshot" of="/mnt/backups/lvm-vg-${lv}_backup-$(date +%Y%m%d).dd" bs=64k
# remove the snapshot
lvremove -f "/dev/vg/$snapshot"
That method is certainly not perfect, as you are still taking the backup with a mounted FS in place and the snapshot size must be adapted to the use-case, but in the end, this is a method to create backups which are complete (independent of FS), easily accessible (e.g. mount via loop device) and do not interrupt the service.Generally when I set up a computer, I use automatic configuration tools, and any sort of data store will have its own automatic backup and restore solution — or even better, won’t be on that server at all.
In the case of a home computer, I’ll install everything using ansible and homebrew (or the App Store or steam) and backup documents and photos to Dropbox and/or iCloud or s3, and all my code and dotfiles of course are on github.
I’ve spilled a soda on my laptop and been back up and running with a replacement in 10 or 15 minutes.
My whole house could burn down and in terms of data, I’d be at best mildly inconvenienced.
For people that don’t trust cloud, it’s way more likely that your physical copy is going to be lost or ruined than that multiple cloud services are going to lose your data.
10/10, would install again.
I don't have backups myself, but I have 5 copies of my data at 3 different locations + snapshots everywhere: desktop, laptop (linux, BTRFS), NAS at home (FreeBSD, ZFS); off-site dedicated server at OVH (FreeBSD, ZFS); at school (FreeBSD, ZFS/netapp, managed by school IT).
Syncthing [0] + whatever handles local snapshots [1,2] work wonders [3]
[1] https://github.com/zfsnap/zfsnap
I've lost data with the Intel onboard raid, several times, so I refused to use the Intel raid implementations; if you are going to utilize RAID, I've learned the hard way to pay for a proper RAID card and cage (or use the tried and true software raid setups).
If you loose the RAID array, take a break before you start fixing it. A client had their subversion repositories with over a decade of work on a Windows server, depending on a RAID setup. It fell down and turns out they didn't have any good backups. The IT guy tried to fix it and nuked the super block and did who knows what else.
They ended up packing up and sending the entire server to some firm that specialized in data recovery and spending a small fortune.
It doesn't matter how reliable your backup is. If you can't restore it you've done nothing. So far I've managed to:
- Backup encrypted content without saving the decryption key
- Backup everything to a remote server without backing up the ssh key to that server
So do yourself a favor, don't just backup. Try restoring some of your content from time to time (or at least imagine how you would do it).
Another scenario I've been recently thinking is, if the house burns down tonight, and you lose everything, phones/computers/etc., would you be able to get your data back?
EDIT: formatting
The bright side is, you'll never make this mistake again. The downside is, for me at least, I end up with triple+ backups of everything..
It was a good read! Thanks for taking your time to write it up.
If that is untenable, mdadm has an option --zero-superblock, lvm has pvremove, and wipefs will remove file system signatures, which you can run before changing the storage layout.
Do you have an example (or rather source) of a RAID metadata outside of the first/last 1MiB of the drive?
I also rsync to an external drive bay that I swap out weekly.
At the end of the day, every single machine in my network of things, including my desktop can light itself on fire - and I don't care. If my office or house is bombed, I don't care. Now - if my city were bombed.. I'd argue that I probably have other fish to fry.
Technically this is a booby trap, not a time bomb. ;-)
https://news.ycombinator.com/item?id=18541493
Should open source software fill up with quirks for the mistakes of other software? This looks like a bug report to other RAID software that should clear up any obsolete flags and enforce a sane condition of the RAID.
# dmsetup remove_all
# dmraid --activate y --format isw
Is that y a "yes" flag !? eg answer Yes on any question.In my experience, working with anything beside simple mirrors in a RAID config will eventually result in loss of data, and that most data loss is because of human failure rather then disk failure.
Dead easy, 10 min, can't fail alternative for home projects: use a VM instead, switch off the VM, copy the disk file, switch the VM back on. Use NVMe SSDs and depending on the size, you can probably do an automated daily backup with minimal downtime.
Do we reduce failure rates to "acceptable" at the expense of neglecting recovery?
That's quite a claim. Was there really no way to recover anything, even with testdisk or photorec?
I had a professor back in 2008 that hammered this into my brain and I share it every chance I get.
i'm basically asking what the impetus was for "just going for it" which is probably going to be a reason that other people share...
Personally I prefer full backups - they feel more complete.
But also, there are different styles of incremental. Some are dodgier than others.
My projects so far ranged from 1-100GB so a full daily backup is quite quick done and its more easy to recover compared to applying incremental backups individually.
Well, always choose the right tool for the task.
At least one style of incremental backups can be restored in a single copy - I use rsync to create hardlinks to the previous day if the file is unchanged (based on modified time). This results in around 400M of additional disk usage per day due to changes, for around 90G of data in my home directory.