Intuitively I would think that the amount of extra writes is pretty low, even if you snapshot very frequently.
But scientific measurements would be nice.
I used to do snapshots every minute, every hour and every day with ZFS on some servers I administered. I’d purge the minute snapshots after 60 minutes. And I had cron jobs on other machines to backup the hourly and daily snapshots. I had it set up so that hourly snapshots were kept for something like 72 hours. And the daily snapshots were kept forever.
The idea with the every minute snapshots being that they were for undoing manually made mistakes during SQL migrations etc.
It worked well for me.
I still use ZFS on my FreeBSD servers. But at the moment my projects are low traffic and the data only changes in important ways some rare times. So with my current personal servers I manually snapshot about once a week and manually trigger a backup of that from another server.
Another thing I’ve changed is that now I only snapshot the parts of the file system where I store PostgreSQL databases and other application data. I no longer care so much about snapshotting the operating system data and such. If I have a serious hardware malfunction I will do a fresh install of the OS, and I have a log of what important config values are used and so on, that my backup scripts copy when I run them, without copying all of the other things.
Unless you're a custodian of some secret society's files!
But it wasn’t. It was very useful in fact.
But if you stop the running DBMS before you restore the previous version on disk, and then start the DBMS again it should be able to continue from there.
After all that’s one of the key selling points of a bonafide DBMS like PostgreSQL, that it’s supposedly very good at ensuring that the data on disk is always consistent, so that when your host computer suddenly stops running at any point in time (power outage, kernel crash, etc) the data on disk is never corrupted.
If data is corrupted by restoring from a random point in time in the past, that should be considered a serious bug in PostgreSQL.
You can of course the up in a dirty state where some new transaction was started but not committed, but any non-toy database should be able to recover from that. (You'll lose the transaction of course)
It is correct though, that trying to do it with something like 'cp' on a live database will most likely be corrupt.
as long as the filesystem supports some kind of journaling (aka it’s not ancient) and the database is acid compliant, there shouldn’t be any major issues beyond a slow startup.
but keep in mind that you may lose acid compliance by fiddling with the disk flushing configuration, which is a common trick to raise write speed on write choked databases. if that’s the case you may lose some transactions
your database documentation should have this information.
Yep :)
At said company where I was using ZFS snapshots for the servers I administered, I additionally had a nightly cron job to dump the db using the tools that shipped with the DBMS. Just in case something with the ZFS snapshots fricked itself :)
I did a daily dump of my prod database to my local workstation, needed it for prod corruption issue because ops was only doing weeklies.
The snapshot doesn't write much, and both SSDs and ZFS are copy on write. Which means the cost of writing after a snapshot is the same as before the snapshot.
On the other hand context is missing. Both SSDs and ZFS don't like being full or even close to full. The working set was ~650GB, of the drive was 1TB, then those snapshots could have easily made the drive over 90% full. This could have made ZFS unhappy all by itself.
I didn't understand this, could you please clarify?
If there was no snapshot, there would be only one write operation, the actual write. However, with snapshot in place, in addition to actual write, there is a copy operation which copies the original data and writes to snapshot location. So, there should be two write operations (actual + copy).
ZFS really deeply assumes that, when a region is in use, it will not change until it's no longer in use anywhere, and it also won't reuse things you just freed for a certain number of txgs afterward to let you get away with having to roll back a couple txgs in case of dire problems without excitement. (Since having enough writes will cause more txgs to happen faster, this isn't an issue people run into with being unable to use newly free space in practice.)
Also in practice, defining what "sequential" means with multiple disks in nontrivial topologies becomes...exciting anyway, and for writes, you only care that things are relatively, not absolutely, sequential for spinning media, and on reads, prefetch is going to notice you doing heavily sequential IO and queue things up anyway. (IMO)
If you like, you could go check on your configurations, what the DVAs for the different data blocks in your VM images are - something like zdb -dbdbdbdbdbdb [dataset] [object id, which you can get from the "inode number" of the file, or if it's a zvol, I think it's always just 1 that all the data you think of as the "disk" goes in...]
You'll almost certainly find that the regions that changed more than a couple txgs apart (the "birth=" value is the logical/physical txg the record was created) are mostly not remotely sequential.
(Nit - the two exceptions that come to mind are, the uberblocks are basically a fixed position on disk relative to the disk's size, and a fixed size, and you get [fixed size]/[minimum allocation size] of them in a ring buffer, basically, before you overwrite the oldest one, and that happens by just overwriting it, since it's technically not in use any more, someone just might want to roll back to it in a "This Should Never Happen(tm)" case...or the newly added feature of corrective send/recv, to let you feed ZFS a send stream of an "intact" copy of something that had an uncorrectable data error and have it scribble over the mangled copy with the fixed one in-place, assuming it passes the checksums.)
Edit: after reviewing a few benchmarks, the outcome seems to be - even on SSD, make sure you actually want the zfs features, because ext4 will be a lot faster.
Because there are various mitigations and configurations involved if you're trying to do lots of small random IO for ZFS, and I've not heard people giving the advice of "just don't" in most use cases.
IMHO, zfs is a clear win for durable storage for documents and personal media. It's not a clear win for ephermeral storage for a messaging service or a CDN. If you don't mind running multiple filesystems, zfs probably makes sense for your OS and application software even if your application data should be on a different filesystem.
* I could be wrong, I asked for some help from not the most reliable sources. Happy to be corrected. Still, if my estimate is higher than actual and yet still unlikely to affect drive longevity, it may be moot.
Official quoted specification for SN850 is 600 TBW of write endurance, likely after derating for obvious warranty implications. Incidentally, 2500TB is also a typical endurance figure for many SSDs in this market. Overall, to me, sounds not entirely impossible.
I kind of wonder what's the controller says in SMART data, if still alive. On Linux the command is `apt install smartmontools; smartctl -s on /dev/sda; smartctl -A /dev/sda`, and it shall print out a table[4]. On Windows, just install CrystalDiskInfo[3].
1: https://news.ycombinator.com/item?id=29165202
2: DWPD: drive writes per day, TBW: Total Bytes Written - in terabytes
3: https://crystalmark.info/en/software/crystaldiskinfo/
4: Note that "Pre-fail" means the value is supposed to change when about to fail and "Old_age" means the value is supposed to indicate age, NOT "this is bad and about to fail" and "this drive is old". It always says all Pre-fail and Old_age. Someone should have changed it to "somewhat_boolean" and "life_remain" long time ago in my opinion.
Your reference says erase block size. That is not minimum write size. Sorry but you are clueless (and careless).