I only lost 10 minutes of data, thanks to ZFS
mastodon.social
mastodon.social
https://petapixel.com/2023/08/08/sandisk-portable-ssds-are-f...
https://www.theverge.com/22291828/sandisk-extreme-pro-portab...
https://news.ycombinator.com/item?id=37042587
https://www.theverge.com/23837513/western-digital-sandisk-ss...
Both HGST drives too. A very sad day.
Thankfully I had been regularly zfs sending my contents to another site and lost very little data.
ZFS is rad.
* https://www.cs.cmu.edu/~garth/RAIDpaper/Patterson88.pdf
* https://www.computerhistory.org/storageengine/u-c-berkeley-p...
Sounds like it was the backup which saved the day here, not the raid array.
And he could use his system without replacing any drive and he would have had a much quicker recovery and could have used the system during rebuild.
Now, raid on a laptop is mostly reserved to bigger units. And authors setup is awesome as well, but single drive failure is one of the most common issues (especially where you make use of snapshots) you need your backup and that is exactly what raid solves.
So try to do both :)
2 drives in a mirrored zpool, this is the equivalent of RAID1.
* They were manufactured in the same batch, maybe even one right after another on the same line.
* As they were transported from manufacturer to OEM to you they were exposed the the same environmental conditions, right down to vibrations, humidity, and ambient EM environment.
* As you use them, they continue to be exposed to the same environmental conditions, including power supply fluctuations and power inductively coupled into places it doesn't belong.
* They see the same usage patterns. Depending on the RAID specifics, that might be right down to seeing the same disk locations seeing the same read and write volume.
Its then not surprising if they fail at about the same time.
The last machine I put together that I wanted to have high availability, I intentionally bought two different brand drives to put in the mirror to maximize the likelihood that they fail at very different times.
Many years ago (c. 2003) the group I was working in inherited a massive 6U storage server with an insane number of 10k SCSI (it was before SAS was a thing) drives. We named it "hurricane" for the sound it made. After a few weeks of using it, the first drive failed. It rebuilt to a hot spare and we ordered and eventually installed a replacement. A few weeks later, another drive failed, and this time before it could finish rebuilding, two more drives in the RAID failed and its contents lost (but we had a good backup). We never used it again. For a while I used it as a coffee table, but then someone convinced me that was too tacky, and it got ewasted.
When setting up a new machine with zfs I intentionally buy drives from as many different brands and models as possible to spread the manufacturing defect risk.
It is, but in a different way. It is a testament to the depth and precision of manufacturing process control, that two insanely complex machines will behave nearly identically for years, up to the point of failing at about the same time, if they've been made in the same batch, and exposed to about the same environment and usage patterns over those years. You'd expect any number of random factors to cause one drive fail way before the other, but no - not only there is very little variation between drives in a batch, tiny variations in usage are damped down instead of amplified.
It truly is amazing.
Your model isn't exactly bad, but there is an assumption being made that you haven't accounted for. Which to be fair, is frequently not stated. The assumption is that the drives defects are independent of one another. This is a poor assumption when manufactured back to back.
I wasn’t using them in a RAID configuration, but they were attached to a raid controller.
I thought it might be the controller but a year on I've had no further issues. Sometimes drives do just go like lightbulbs.
Take it as a good impetus to catalog your data and find those extra replication options. Data protection does cost you a bit to do it "fully", but since I've worked on a backup solution before in a client facing way, trust me when I tell you that I've seen rather large businesses (a few you might even know as a household name) who have less consideration for their data than you've expressed in your 3 sentences :)
So just figure out which of your data _truly_ needs to survive at all costs, get a solid setup with personally owned storage for the backups in combination with cloud storage, and you're probably fine.
Alternatively, consider just using a file system that does periodic integrity checks. I know ReFS has integrity streams, and I am pretty sure XFS has something almost exactly the same or better. It won't prevent corruption, but it will give you something you can monitor for when the filesystem reports an issue.
Some combination like this should work.
Similarly, you might be able to come up with a fast trick using stat; with some quick testing on a dummy file in MacOS' ZFS shell, you can do something like:
stat -f %m somedir/*
and compare the resulting value by passing it to sum or something. I am not super familiar with stat in general, so likely I am missing elements that make this unreliable, but I'd consider looking into it further unless someone tells me it's 100% the wrong direction and explains why.
ZFS has checksums, and the 'zpool scrub' command tells it to verify those (on all copies of your data, if you're using RAID).
it has become non economical / practical for me to backup everything
Using matched drives seems to be a very bad idea for mirrors. I'll probably replace them with two different brands.
I also have a matched pair of HGST SAS Helium drives in the same backplane so hopefully I can catch those before they fail too if they're going to go at once, I _do_ have data on those.
Normally, failures come from some amount of non-repeatability or randomness that the systems weren't robust to.
The drive industry is special (in a bad way) in that they can exactly reproduce their flaws, and most people's intuition isn't prepared for that.
Since mine had ~43000 hours, they didn't fail prematurely, they just aged out, and since they appear to have been built pretty well, they both aged out at the same time. Annoying for a ZFS mirror, but indicates good quality control in my opinion.
> Google assumes you’re looking for product pages when you search for things like “4TB SanDisk SSD,” so news stories like ours and Ars Technica’s appear far down search results.
Well, I've been test-driving the paid Kagi search engine, and this was an excellent opportunity to see if a different class of web search could produce different results...
Sadly I'm afraid when all the web is flooded by praising articles, a different set of prioritization rules would still struggle to show different results:
* First 5-6 results are from shops.
* Then come some reviews, from: easeus [1], techpowerup [2], anandtech [3], consumerreviews [4]. None of them contain the word "fail".
* Lastly, and this is the only actual improvement from google, there are some relevant search suggestions such as "sandisk 4tb ssd failure" and "sandisk 4tb ssd problems". Difficult to see (they are at the bottom of the page), but at least better than Google (where the word "fail" doesn't appear at all in the first results page).
[1]: https://www.easeus.com/knowledge-center/sandisk-4tb-extreme-...
[2]: https://www.techpowerup.com/review/sandisk-ultra-3d-4-tb-ssd...
[3]: https://www.anandtech.com/show/16892/sandisk-extreme-pro-cru...
[4]: https://consumerreviews.store/sandisk-4tb-extreme-portable-s...
Now btrfs isn't ZFS, but it has some feature parity and perhaps the "poor man's ZFS". It's also much more reasonable to run on certain OS, due to the licensing, packaging, and in-kernel status of ZFS being kind of weird.
One memorable time I was encouraged to use ZFS was when I mentioned to the Linux User's Group that I'd had to pull the power cord to reboot my computer, and I was roundly scorned for this foolish maneuver. But you may change your mind about the wisdom of doing either one when you consider that the system in question was a Raspberry Pi. Heh.
I still love Linux and I'd use it for any given server or Raspi if that were part of my job. I do use it daily in my job, but to a very minimal extent.
ZFS looks so cool! Unfortunately when I eventually get a NAS I doubt I'll want to pay for anything that can run it, so I suspect I'll just be doing RAID and ext4.
I always stayed away from BTRFS, because every few months I'd see a "BTRFS destroyed my data" post, followed by an argument about if it was BTRFSes fault. I see them less now, perhaps it's time to revisit?
I personally run my nas with Ubuntu and ZFS and love it.
But thanks for bringing this to my attention. I had missed the changes in 23.04.
FreeBSD won't break your wallet.
It'll run fine but more sadly on older Pis running 32-bit kernels, since it does a looooooooot of 64-bit and wider operations, so you pay a nasty tax on that on 32-bit things. (Though the virtual address space limits might actually be sadder than the 64-bit operation penalty there, really...)
>Misinformation has been circulated on the FreeNAS forums that ZFS data integrity features are somehow worse than those of other filesystems when ECC RAM is not used. That has been thoroughly debunked. All software needs ECC RAM for reliable operation and ZFS is no different from any other filesystem in that regard
Of course if you have good backups, you can use whatever and not really worry too much about it.
For the most part though, sticking to standalone or mirrored disks is pretty rock solid and has been for a long time. Ditto for subvolumes, snapshots and, send/receive. My laptop has been snapshotted and sent from one piece of hardware to the next for many years now.
That said, I'm with you on the backups. Anyone who uses btrfs and doesn't have a rock solid backups is a mad man.
Ext4 seems to handle that scenario better. I can't think of a single instance of filesystem corruption that fsck couldn't fix, and some of those VMs have probably been abruptly terminated at least a hundred times over the years.
I only ever had these issues with SD cards. I quickly switched to running my pi off a USB external SSD, and haven’t had any problems since then. Now when the power goes out, it boots back up properly and all my services start. All of this on ext3 I think?
Planning to redo things for ZFS at some point, but haven’t gotten around to it yet.
Our $job dashboards used to nuke an SD card every couple weeks/months, but since the move to logs-in-RAM we've been running the same SDs for years.
DIY via {fs,journalctl} config , or using https://github.com/azlux/log2ram
Also, mount the SD with the `noatime` flag of course: https://wiki.archlinux.org/title/Ext4#Disabling_access_time_...
From what I hear Home Assistant is still not the easiest if you want to run for years on a card, not sure if that's fixed now, but it's one of the big factors blocking me from moving to HA.
Note: only a guide, not scripted yet. See footnote for why :)
https://github.com/EternityForest/KaithemAutomation/blob/dev...
https://www.digikey.ca/en/products/filter/memory-cards/501?s... is the brand I used to use, iirc
Note that this doesn't have the Apache logfile hack so Apache probably won't run if you try this and don't add something to make it's fussy logfile.
https://github.com/EternityForest/KaithemAutomation/blob/dev...
Seems like people say it is more CPU heavy than EXT4, so unless it's way more reliable, would it really be the best choice on a pi/router/subGHz commercial NAS chip?
Why do you need/want RAID ?
RAID is for availability, but is your data really that important that you cannot wait for a restore ? Most people would be much better off using that 1..N parity drives as versioned backups instead of running RAID.
I've run NAS boxes for years, but these days i'm only using single drives.
My setup these days consists of laptops that synchronizes data (encrypted) to the cloud, and a small ARM machine that synchronizes cloud contents locally, and makes a versioned backup to a single drive as well as a versioned cloud backup.
As for cost, it's cheaper (for me) to store my data in the cloud, than the cost of electricity required to run my NAS.
Isn't RAID parity slightly more space efficient than versioned backups? Or is there a better way to do redundancy that doesn't involve just replicating entire files to multiple disks? Or some kind of automated manager that puts each individual file on N different disks out of M?
I mostly do embedded so reliable data storage isn't generally something I deal with, we usually leave that to the cloud or to the user, and I'm not quite familiar with what's out there.
It depends on your storage array. The more drives, the more space efficient RAID becomes, but RAID is still only a single copy of your data.
>Or is there a better way to do redundancy that doesn't involve just replicating entire files to multiple disks
Most of the industry is using erasure coding these days (https://blog.min.io/erasure-coding/) which allows for spreading your parity and data across multiple sites. Erasure coding usually runs a layer above the filesystem, as opposed to RAID which typically runs below the filesystem (Snapraid, Mergerfs and others excluded).
My personal "backup vault" is a Raspberry Pi 4 with a single 4TB external drive attached. The RPi runs Minio, and all backups are done through the S3 interface or SFTP/SMB. It is not the fastest box in the world, but it backs up (incremental) ~2TB in 30 minutes, which is "fast enough".
It consumes on average 4W, which means even with worst case electricity prices of €1/kWh (which we saw last winter), it costs less than €3/month.
For comparison, my NAS consumed around 50W, and at €1/kWh, that would cost €37/month in electricity alone, and then you need to add the cost of the actual hardware itself.
I switched off the NAS, and purchased ~10TB of cloud storage (main storage and backup storage at two different locations) for €20/month, and keep sensitive stuff encrypted with Cryptomator.
Btrfs has 3-copy and 4-copy RAID1 (profiles raid1c3 and raid1c4), doesn't ZFS have something similar?
My point was that even if you have 4 copies of your data, you still only have a single machine where your data is stored, and you're essentially just one flood/lightning strike/house fire/burglary away from all of it being gone. Or one bad power supply away from 4 dead drives.
With versioned backups, you have higher latency on restoring data in case a disk dies, but your data is also safer.
As i initially stated, RAID is for availability. It is great for making sure that data is available 24/7, but that is rarely what the average home user needs. Most home users access their data infrequently, and would be perfectly fine waiting a couple of hours while restoring data from a backup.
how much data are we talking about?
There is a turning point somewhere after 15TB where the cloud becomes a lot more expensive than local storage, but at 10TB it’s hard to do in a reliable manner locally.
And it doesn’t really matter how you do it. If you’re the DIY type, and buy a second hand server, or a small machine, or just run out and buy the latest Synology box, you’re still paying more for storing <15TB at home in RAID. One will be more expensive in electricity, the other in purchase price.
The average price of power during a year here is around €0.45/year, so a box drawing 50W (not unlikely with a decade old processor and two disks) will use 438 kWh during a year, meaning it would cost €197/year to keep it powered.
Add to that the ~€150 x 2 for new drives, and you end up with €227/year, or €19/month.
But of course, if you, as you stated, had the box running anyway, the math works out differently as you're essentially splitting the cost with whatever purpose the box already fulfills.
I ran ZFS on a Raspberry Pi 4 with 8GB of RAM just fine (under debian arm64), and I've used ZFS on a machine with 4GB of RAM for receiving snapshots.
In terms of CPU and RAM, most NASs now are perfectly capable of running it (but do also consider speccing out a small form factor PC with a case with lots of drive bays vs prebuilt NAS - you might be surprised at what you can get for a similar cost).
Tell me about tmpfs: where do you use it?
Eventually I'll probably move it to a standalone thing since there seems to be a whole lot of interest in the SD protection feature, but it's missing a few of hacks for programs I don't really use anymore, like the apache logfile thing, just because I got tired of maintaining stuff that didn't have much interest.
zvols are so much better for that
The internet is full of problems lately about data loss and longevity issues.
I remember 20 years ago HDDs were not meant for eternity either, but they definitely outlived the usefulness of the computer that they were bought with...
We're hitting scaling limits. Exponential growth is slowing.
Meanwhile today I'm helping admin a ZFS server with 20+ drives and drives have about a 4-5 year lifespan.
> but they definitely outlived the usefulness of the computer that they were bought with
Computers were also much more quickly obsoleted back then. When today a 6 year old computer is totally useable, back then you really felt it if your machine was just 3 years old.
Was at the Aussie Tribes 2 launch LAN and there was a guy who had one die on him.
At that time in the LAN scene there would always be someone who had a Deskstar die ... You could hear the clicking over the noise of the LAN.
I realised back then, I can only trust Seagate.
I have even older memories of problematic IBM drives. During the early 90s the shop I briefly worked with, found a supplier for IBM SCSI drives at a very convenient price, so they ordered a good lot of them. They worked great on PCs, but some of us also had Amiga machines and of course would love to benefit from the offer. So we tried one, but it didn't work; then another, and another; nothing, they were normal SCSI drives but refused to work on any Amiga with a SCSI controller, although any other drive would work in there. In the end we abandoned all hopes and took the drives for a reformat to be used on PCs, but... they were all dead. Completely, not even detectable by any controller; the mere connection to an Amiga SCSI controller destroyed them instantly. We never discovered where the problem was; those drives worked perfectly on all PCs, while we could install any other drive on every Amiga and expect it to work, but no way to put those in an Amiga and expect it to survive. Good old times indeed:)
I think it was about 10 years ago some Seagate drives had higher than industry failure rates. Iirc one of their factories was producing drives that failed much more frequently than others (there might have also been something to do with the platter counts/model)
Anecdotally I remember HDDs failing sometimes for me and my friends/relatives back in the day. Now it barely happens for SSDs. Hell, even supposedly problematic old Intel from late 2000s still works fine in the same old MacBook I gave to my mother after using it for years.
I wonder what it the actual data regarding this.
0: A couple exceptions: Two WD Red (before SMR) which I took out from my old NAS to put bigger drives in place, and put in a drawer while they were still perfectly healthy. After like 2.5 years in their anti static bags and normal conditions, no excessive heat, no moisture, no magnetic fields etc, I took them out because I needed a spare disk and checked them: both were not working, one completely dead and the other barely recognizable but unreadable. The first didn't even show up once connected; I tried to clean all contacts, including the pins on the controller pcb to no avail, and eventually had to ditch it; the 2nd one could be reused only after a full reformat; no way to recover old data, not even using testdisk. I never experienced nor expected anything like that, and frankly it worries me quite a lot.
Would be cool if it could just use e.g. S3 for storage.
Between this and NixOS, I can provision a new identical laptop in about 10 minutes.
I recently added off-site replication as well, so even if I get completely devastatingly mugged, there's still about zero chance of serious data loss.
Zrepl is absolutely brilliant software. Easy to run with, but incredibly sophisticated and powerful if you need all the knobs. I can't praise it enough.
Intuitively I would think that the amount of extra writes is pretty low, even if you snapshot very frequently.
But scientific measurements would be nice.
I used to do snapshots every minute, every hour and every day with ZFS on some servers I administered. I’d purge the minute snapshots after 60 minutes. And I had cron jobs on other machines to backup the hourly and daily snapshots. I had it set up so that hourly snapshots were kept for something like 72 hours. And the daily snapshots were kept forever.
The idea with the every minute snapshots being that they were for undoing manually made mistakes during SQL migrations etc.
It worked well for me.
I still use ZFS on my FreeBSD servers. But at the moment my projects are low traffic and the data only changes in important ways some rare times. So with my current personal servers I manually snapshot about once a week and manually trigger a backup of that from another server.
Another thing I’ve changed is that now I only snapshot the parts of the file system where I store PostgreSQL databases and other application data. I no longer care so much about snapshotting the operating system data and such. If I have a serious hardware malfunction I will do a fresh install of the OS, and I have a log of what important config values are used and so on, that my backup scripts copy when I run them, without copying all of the other things.
Unless you're a custodian of some secret society's files!
But it wasn’t. It was very useful in fact.
But if you stop the running DBMS before you restore the previous version on disk, and then start the DBMS again it should be able to continue from there.
After all that’s one of the key selling points of a bonafide DBMS like PostgreSQL, that it’s supposedly very good at ensuring that the data on disk is always consistent, so that when your host computer suddenly stops running at any point in time (power outage, kernel crash, etc) the data on disk is never corrupted.
If data is corrupted by restoring from a random point in time in the past, that should be considered a serious bug in PostgreSQL.
You can of course the up in a dirty state where some new transaction was started but not committed, but any non-toy database should be able to recover from that. (You'll lose the transaction of course)
It is correct though, that trying to do it with something like 'cp' on a live database will most likely be corrupt.
as long as the filesystem supports some kind of journaling (aka it’s not ancient) and the database is acid compliant, there shouldn’t be any major issues beyond a slow startup.
but keep in mind that you may lose acid compliance by fiddling with the disk flushing configuration, which is a common trick to raise write speed on write choked databases. if that’s the case you may lose some transactions
your database documentation should have this information.
Yep :)
At said company where I was using ZFS snapshots for the servers I administered, I additionally had a nightly cron job to dump the db using the tools that shipped with the DBMS. Just in case something with the ZFS snapshots fricked itself :)
I did a daily dump of my prod database to my local workstation, needed it for prod corruption issue because ops was only doing weeklies.
The snapshot doesn't write much, and both SSDs and ZFS are copy on write. Which means the cost of writing after a snapshot is the same as before the snapshot.
On the other hand context is missing. Both SSDs and ZFS don't like being full or even close to full. The working set was ~650GB, of the drive was 1TB, then those snapshots could have easily made the drive over 90% full. This could have made ZFS unhappy all by itself.
I didn't understand this, could you please clarify?
If there was no snapshot, there would be only one write operation, the actual write. However, with snapshot in place, in addition to actual write, there is a copy operation which copies the original data and writes to snapshot location. So, there should be two write operations (actual + copy).
ZFS really deeply assumes that, when a region is in use, it will not change until it's no longer in use anywhere, and it also won't reuse things you just freed for a certain number of txgs afterward to let you get away with having to roll back a couple txgs in case of dire problems without excitement. (Since having enough writes will cause more txgs to happen faster, this isn't an issue people run into with being unable to use newly free space in practice.)
Also in practice, defining what "sequential" means with multiple disks in nontrivial topologies becomes...exciting anyway, and for writes, you only care that things are relatively, not absolutely, sequential for spinning media, and on reads, prefetch is going to notice you doing heavily sequential IO and queue things up anyway. (IMO)
If you like, you could go check on your configurations, what the DVAs for the different data blocks in your VM images are - something like zdb -dbdbdbdbdbdb [dataset] [object id, which you can get from the "inode number" of the file, or if it's a zvol, I think it's always just 1 that all the data you think of as the "disk" goes in...]
You'll almost certainly find that the regions that changed more than a couple txgs apart (the "birth=" value is the logical/physical txg the record was created) are mostly not remotely sequential.
(Nit - the two exceptions that come to mind are, the uberblocks are basically a fixed position on disk relative to the disk's size, and a fixed size, and you get [fixed size]/[minimum allocation size] of them in a ring buffer, basically, before you overwrite the oldest one, and that happens by just overwriting it, since it's technically not in use any more, someone just might want to roll back to it in a "This Should Never Happen(tm)" case...or the newly added feature of corrective send/recv, to let you feed ZFS a send stream of an "intact" copy of something that had an uncorrectable data error and have it scribble over the mangled copy with the fixed one in-place, assuming it passes the checksums.)
Edit: after reviewing a few benchmarks, the outcome seems to be - even on SSD, make sure you actually want the zfs features, because ext4 will be a lot faster.
Because there are various mitigations and configurations involved if you're trying to do lots of small random IO for ZFS, and I've not heard people giving the advice of "just don't" in most use cases.
IMHO, zfs is a clear win for durable storage for documents and personal media. It's not a clear win for ephermeral storage for a messaging service or a CDN. If you don't mind running multiple filesystems, zfs probably makes sense for your OS and application software even if your application data should be on a different filesystem.
* I could be wrong, I asked for some help from not the most reliable sources. Happy to be corrected. Still, if my estimate is higher than actual and yet still unlikely to affect drive longevity, it may be moot.
Official quoted specification for SN850 is 600 TBW of write endurance, likely after derating for obvious warranty implications. Incidentally, 2500TB is also a typical endurance figure for many SSDs in this market. Overall, to me, sounds not entirely impossible.
I kind of wonder what's the controller says in SMART data, if still alive. On Linux the command is `apt install smartmontools; smartctl -s on /dev/sda; smartctl -A /dev/sda`, and it shall print out a table[4]. On Windows, just install CrystalDiskInfo[3].
1: https://news.ycombinator.com/item?id=29165202
2: DWPD: drive writes per day, TBW: Total Bytes Written - in terabytes
3: https://crystalmark.info/en/software/crystaldiskinfo/
4: Note that "Pre-fail" means the value is supposed to change when about to fail and "Old_age" means the value is supposed to indicate age, NOT "this is bad and about to fail" and "this drive is old". It always says all Pre-fail and Old_age. Someone should have changed it to "somewhat_boolean" and "life_remain" long time ago in my opinion.
Your reference says erase block size. That is not minimum write size. Sorry but you are clueless (and careless).
That works. TM restores can be quite slow, but almost all my important data is in Git (and hosted storage), so it’s not really been an issue. I just use TM every now and then, if I have a single file I want to backtrack.
I also have one of the notorious[0] SanDisk drives. I don’t use it for anything important. It just has some game storage. Since I’m a Mac user, games aren’t really much of a factor for me, and I won’t cry, if they croak.
[0] https://arstechnica.com/gadgets/2023/08/sandisk-extreme-ssds...
Time Machine really doesn't like using remote disks that aren't official Apple gear.
- https://eclecticlight.co/2021/03/11/time-machine-to-apfs-und...
- https://eclecticlight.co/2021/04/16/time-machine-to-apfs-usi...
Not gonna lie, it was pretty terrifying until I had my first confirmation I could decrypt the data.
Note: no bespoke backup method should be assumed functional unless you actually periodically check that you can restore data from it.
You can't restore to production obviously, as aside from the downtime if the test fails you've just destroyed your good copy and proved the other copy is also bad.
About the best I can think of is restoring a small part of the set as a sample, which isn't really testing the whole thing.
Any cool things?
Until then per-directory data replicas is the killer feature for me (Music has 3, Documents has 5, Downloads has 1). Something to be very excited about with full compression and encryption.
I don't use multiple replicas, but I use that to tailor my backups per directory. ~/documents is snapshotted and backed up on the regular, with long-lived snapshots. Code is snapshotted regularly, but snapshots don't live too long, and they're not shipped to a different drive. I don't care for ~/tmp so no snapshots.
---
edit: the cost, besides having to actually create the file systems, is that moving data between them isn't instant.
No more the case after block cloning support goes production: https://github.com/openzfs/zfs/pull/13392
https://www.phoronix.com/news/Linux-Torvalds-Bcachefs-Review
(list most likely incomplete)
Sounds like the process was a bit less tested and documented than optimal. For a home system or personal desktop that's not super unusual.
You don't want to be working out your restore procedure on the fly for production servers though. ;)
[1]: https://github.com/zfsonlinux/zfs-auto-snapshot
[2]: https://pilabor.com/series/proxmox/restore-virtual-machine-v...
Or put it more directly, full disk backups are a great way to get RTO down.
I've lost data before, but it felt terrible to lose my context and working memory. While I make sure the most important stuff is in git, there's a bunch of momentum and working memory in my bash history and system configuration. It's also nice to not have to think very hard about a patchwork of backup plans.
It's nice to get a fresh start every now and then, but not under duress. I was in the middle of a multi day project and was gonna lose time either way. It was real nice to boot back into a machine that felt like home.
If for any reason any one of these systems either is destroyed or is no longer master, picking up from where I left off is as simple as picking another system up and marking it "master". No restore process, no changes, nothing at all, and it picks up from exactly where I left off when I was working on the other system. Means I can just grab my EDC laptop and stuff it in a bag not knowing how long I'll be out or where I'll be going and also know that it will be completely up to date with my datasets, or I can grab my desktop replacement laptop and its enormous external disk if I am going to be on a different continent for an extended period of time and want full geographic dataset locality. At no point in time does any of the above require the manual running of any process or replication or anything like that.
Reprovision and restore would take a whole lot longer than this, wouldn't give the abilities that it provides, and the above is only possible because of zfs snapshot replication.
I also use a USB C external SSD that is a member of a ZFS mirror and a md raid group that is bootable, so even if my EDC laptop were to spontaneously combust, I could immediately get up and running on any similar laptop with roughly comparable hardware simply by putting that SSD in and booting from it, then adding the SSD on the laptop to the zfs mirror / md raid group.
Do you mean don't let this story fool you to not test your backups? Because the whole point of the story is he was saved by having his backups. (Though you're right that he lucked out by having it work when he hadn't tested it)
It's also scary in general to go from 2 copies to only 1 copy of data. A friend and I have been planning to trade replicas but haven't set it up yet. There's definitely still room for improvement in my setup.
I imagine it is, but I don't know how.
I don't think the performance and drive-lifetime hit of running rsync every 10 minutes would be good.
Zfs should have an edge in terms of atomicity as well, but in practice I'm not sure how much that matters. I _think_ it does matter but isn't perfect (zfs can't trick applications into doing atomic writes if they're not already, but it won't have _another_ worse layer of breaking atomicity like rsync must).
If you read all the files on both sides, every time and don't have more ram than disk, it's going to be a lot of work. If you're just looking at directory entries most of the time, there's a good chance that's all cached and it's no disk load, other than the small changes.
I ageee with you though that atomicity is a big difference, if it matters, and in most cases, it probably doesn't.
Personally, I've mostly stopped doing rsync backups in favor of zfs send, but I've still got one I need to get around to changing. Sanoid/syncoid is pretty decent for less effort snapshotting and syncing snapshots; but I haven't done anything with encrypted datasets. For most of my systems, I'd prefer recovery over security. For the one system in iffy hosting, it runs full disk encryption as a layer below zfs, so it's zfs sends are cleartext, too. (The hosting facility has given me other customer's unwiped disks; better for me to assume my disks won't be wiped)
Yeah that's a good point. I know that rsync is _quite_ clever, but at least any incantations I've ever done it still hits the drives a good amount. I'd ballpark guess a couple of orders of magnitude better than just "cp -r" or something, but still a couple of orders of magnitude worse than zfs snapshots.
Yeah you're 100% right it'll depend on bunch of variables though.
> Sanoid/syncoid is pretty decent for less effor snapshotting and syncing snapshots; but I haven't done anything with encrypted datasets.
I'm not sure I'd recommend it, but I use both directly on encrypted datasets. I have tested recovery a couple times and it works fine, but I've read some cautionary tales too. I _think_ they're all old issues?
My impetus for swapping to a send-based approach from an rsync-based approach of running an incremental (just from the normal non-snapshot filesystem view) then taking a snap on the far side was encountering corrupted encrypted containers. If rsync ran while a container was mounted and being written to, it'd get an inconsistent view of the underlying file and produce a nonsense diff, resulting in an unmountable container in the backup. This doesn't happen with send because it has a consistent view of the blocks it needs to replicate, but as I was writing this I realized that running rsync from zfs's view of the snapshot might get around that. Of course, at that point it's probably easiest to just use send.
I didn't originally have that area as its own zfs filesystem, and it wasn't even originally on zfs, but I moved things around when setting up a new, offsite, backup system... Just haven't gotten around to redoing the old backups. I don't think I'd spend any more time on rsync based backups, given how I use things now; incremental zfs means no need to compare, which makes me have good feelings.
Sure, having to scan a whole tree of files can take a toll on the performance as perceived by other apps trying to use the drive.
zfs dedupe is pretty expensive and doesn't often work out like people might expect...
Sadly for safety's sake directory hardlinks pretty much don't exist, so this doesn't save as much as it could. Apple hacked in an exception for Time Machine so they could get these additional savings.
* People on macOS don't have ZFS, well... maybe they could? See https://github.com/spl/zfs-on-mac
In the end I chose ZFS for the efficiency of snapshots (vs. a full disk scan) and atomicity. Both enable more frequent, smaller syncs, which is perfect for a laptop.
In some ways it seems very attractive as you can just get a low-power SoC (like a Raspberry Pi or a NUC) and hook it up to the external drives.
But there also seems to be many potential pitfalls. Like, how slow will a resilver be over the USB? Might it be unusable/dangerously slow? How reliable is the USB connection? Does it perform to spec or might it cause weird issues? Can you get SMART info over the USB connection? Other issues?
The N in NAS means "network" as in network attached storage so it's not a NAS.
> really curious about how wise such a NAS setup is?
For data use cases like this, USB 3 can be reasonably comparable to Thunderbolt 3, and that connection is generally faster than the media.
This use case seems to be using the external device as a continuous external backup rather than as network attached storage, which is a great use of USB-C dongle SSD enclosures that are the same size or larger than the SSD inside the laptop.
You effectively have mirroring as well, since you have both the internal SSD copy and the external SSD copy, in different makes and forms, unlikely to both fail.
On a related note, I can't think of any drive that has a network interface instead of something like USB, Firewire, Thunderbolt, SCSI, IDE, etc. so how exactly would you define a NAS device?
> For data use cases like this, USB 3 can be reasonably comparable to Thunderbolt 3, and that connection is generally faster than the media.
External HDD enclosures can often contain 4 drives. During a resliver all of these could be heavily accessed. I'm not sure how the single USB 3 connection fares in this scenario. In a normal desktop you'd have four separate SATA connections, and even then resilvering a large RAID setup can take quite some time.
Unfortunately, I haven't had luck with Open Source backup software that uses it (the shadow copy snapshot would fail, the error code would be no help, and finding no resources, I gave up), but commercial software I've used was great. When I was at a big corp, the commercial backup software whose name escapes me at the moment would litterally wait for files to be saved, then do an incremental backup.
As of now, I'm using Veeam on my personal machines, and it runs an incremental backup nightly and saves to an smb share.
At the moment, I use Nextcloud to sync data to my server. It is a more selective approach and Nextcloud is, per se, not a backup solution because not all files can be backed up.. and live-sync is always a half-baked backup solution.
If you want similar features, I think ReFS comes close. AFAICT it's not supported as a boot drive.
I see people use GPU pass-through to play games on Windows VMs. You could probably pass through practically all devices (GPU, sound, network, keyboard, etc.) and this could work well-enough if you don't need the absolute last drop of performance from your CPU and drives. And since KVM supports nested virtualization, you could run WSL2 in the Windows VM.
And if I'm not mistaken, the KVM agent in windows can be told to ask the guest OS to sync the drives, and some Windows applications [0] can even cooperate with this and flush their buffers to disk. You could signal this before creating the ZFS snapshot.
[0] probably not most, but I think MSSQL does.
Anyway, these days, I work 60% of time on my VMs, and the other 40% in WSL. Windows is reduced to a graphical interface, maybe I should simply ditch this last mile, too.
[1]: https://www.reddit.com/r/zfs/comments/yxipyy/anyone_using_op...
It is theoretically possibly to construct a scenario where evil ram does all the exactly right things needed fool ZFS and corrupt your filesystem. Any pearl clutching about this thing which has never happened somehow also ignores that every filesystem is going to get corrupted.
In reality, while ECC memory is always nice to have, it's no more required than any other filesystem. Though personally now that amounts of +32gb are common, I generally prefer error correction/detection over ultimate speed these days. Though ironically ECC memory is actually really nice to overclock, because I can actually just check my logs and prove if my system is actually stable.
There so many actual dangers to your data in comparison that it's laughable. The biggest one being you. Followed by hardware failure, malware, and genuine ZFS bugs. I'd stay far away from raw sends of encrypted datasets in ZFS for a while, there are edge cases that haven't been resolved yet.
Edit Longer article saying the same thing: https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
<sigh> I mean, sure, he recognizes the difference which a lot don't, and I guess yay for zfs here to save him, but this is just irresponsible if you value your data.
which reduces to "I had backups" indeed.
3-2-1 forever!
That wasn't part of the article. It was a single drive failure, so RAID would have done fine.
What he does is run zrep to make a backup. it covers his needs. ZFS snapshot by itself is only transitionally a "backup" for the immediacy of change, it's the least safe form of backup if it remains on the same logical drive structure.
backups also can fail, that zrep can start failing after some os/kernel update without notifying owner. The question is in probabilities of failures, I kinda would trust industrial raid more than some custom made hobby solution.
My SSD boot drive makes me nervous as heck, constantly backing it up.
I think you are very unusual if you don't care about any of this.
Quite possibly.
zfs seems incidental to me. I could have a 10 minute cron job rsyncing changes from ext4 and been just as well off.
But your computer shouldn't run ZFS, that's for the big boys upstairs. Code's too big, it's too hungry.