Why does the integrity of the backup rely on files stored on the computer being backed up? This seems so stupid that I'm sure I'm missing a clue.
But I can't do that without knowing which files are missing.
Reading the explanation in the reddit thread, that's not the impression I got at all.
1. If your computer exploded, your backup integrity would not be compromised
2. If gremlins in your computer did mess with the file, your backups could be compromised. That sounds bad, until you realize that gremlins in your computer could also compromise the executable to do other things that could compromise the backup, (eg. telling the server to delete existing backup data because the retention period has passed or whatever, or simply uploading bad data and waiting for the retention period of 30 days to pass). Moral of the story: if the computer doing the backup can't be trusted to operate correctly, all bets are off.
He doesn't understand the fundamentals of the problem he's solving, and is actively writing code that basically puts tools down and starts shouting "EVERYONE STOP!" in response to expected scenarios.
Most filesystems provide no guarantees at all by default on writes. NTFS journals metadata writes, but not data writes. Append-only files are absolutely expected to be truncated. The application must deal with this either by being insensitive to rollback, or by explicitly requesting a write cache flush! This is very well known to anyone that has ever written any kind of write-ahead-log, database engine, etc... There's a bazillion articles about how this is intended behaviour and no amount of screaming and shouting will make it go away. Learn the storage API guarantees, or STOP writing code that has critical dependencies on storage!
This quote especially seemed childish and immature to me:
> "And this one makes me actively angry, because both Microsoft and Apple will happily throw away portions of your files and not tell you about it"
No, this is NOT Microsoft's or Apple's fault. It is 100% HIS fault for not understanding what's going on. Even if a file flush is correctly requested, consumer HDD and SSD drives are well known to ignore this and keep things in volatile RAM cache to boost their IOPS numbers at the expense of durability. Only "enterprise" drives honour this API properly, and even then there are corner-cases around things like 512/4K sectors and torn writes. Similarly, consumer drives don't have power protection, so partial sector writes are entirely possible and should be expected if they lose power mid-write, or just crash at an inopportune time.
To be more constructive: The correct thing to do is that a log writer must always end log writes with a checksum of some sort. If the log is truncated (for any reason!), then it must recover starting from the last-known-good checksum. Throwing your hands up and saying "NO MORE BACKUPS FOR YOU! EVER!" is not the right response. Retrying backups from the last-known-good point automatically is the correct response.
PS: Some of his other rants are also a lack of understanding of thread safety and/or the lack of ECC RAM in consumer-grade PCs causing random misbehaviour. Heck, you'd also have to deal with bad sectors, corrupt filesystems, and a bunch of other corner cases that just makes this guy scream and shout instead of writing robust code...
PPS: I just realised that the reddit post is by 'CTO and Founder of Backblaze'. Wow. Note to self, do not use Blackbaze for anything, ever. If the CTO is this clueless, then their products can't be trusted.
I'd be more interested in the skill of the developers themselves than the CTO.
Even if you could guarantee some things, refusing to run backups or restores in the face of rollback is just Wrong with a capital W.
Yup, therefore:
| you will always need to be able to figure out the state/diff from scratch.
So if you evaluate a backup system, this needs to be the first thing to check. "what if I backup and accidentally lose my log/summary/sync-state ?"
> This quote especially seemed childish and immature to me:
>
> > "And this one makes me actively angry, because both Microsoft and Apple will happily throw away portions of your files and not tell you about it"
>
> No, this is NOT Microsoft's or Apple's fault. It is 100% HIS fault for not understanding what's going on. Even if a file flush is correctly requested
Several times you seem to jump to the assumption that I don't understand fsync and disk flushing and that is the core issue. You aren't understanding what I'm criticizing. Here is an example of what bothers me:
You take a picture at your wedding, and you store it on your hard drive. You like the photo, it means a lot to you, and you use it as the background for your desktop FOR FIVE YEARS. You have rebooted hundreds of times, and it's always the background for your desktop. Then one day 5 years after your wedding, you reboot your laptop, and it seems to take a little longer to boot, and then after you sign into the laptop half the image you use for your desktop background is scrambled. The middle of it looks like dirty snow. You didn't get any reports of any issues from the OS manufacturer, but now one of your photos is corrupted.
This isn't because the software that wrote the photo 5 years earlier forgot to flush the picture to disk. It just isn't. Behind the scenes, as your laptop was booting from an aging drive, it probably encountered some issue, and it went about fixing the problem as best it could - which I have no problem with. My issue is the drive lost some data, and if the OS manufacturer would let you know this occurred you could take useful actions like order a new drive, prepare a restore from a few weeks ago before that issue occurred, etc.
> No, this is NOT Microsoft's or Apple's fault.
It isn't their fault that the drive is going bad, I agree. Drives go bad, that's why we have backups. My issue is the OS manufacturer try to cover up too much, keep too much hidden from the user, and didn't let the user know data loss has occurred (or might have occurred). And yes, I hold them accountable for not telling customers what is going on. It isn't anybody's "fault" that it occurred, but there is a responsibility to let customers know about it so the customer can take appropriate actions so they don't lose data (or more data).
I try to write incredibly paranoid software. Part of the reason is that is the "average" environment the Backblaze client runs in is more unstable than what most software developers are used to. The whole point of backups is to run when the computer is going sideways, it has bad RAM, it's losing disk sectors, or a customer's cat likes sleeping on the keyboard because it's warm, and the fans are clogged with cat fur. And because the family has teenage children that don't know about computer security problems, they download and install unstable junk from all over the internet because why not? That's the environment my software runs in, and I take my job of trying to protect my customer's data very seriously.
lol