How do I safely store my files?
photostructure.com
photostructure.com
If companies with Hollywood funding consider digital archival too expensive and too fragile... should really tell you something.
Other long-term archives are not digital, either. Germany's Barbarastollen houses thousands of stainless steel barrels with black and white microfilms - about a billion of them.
I'm ignorant of the movie industry. Are you also saying that movies that were 100% digital capture (e.g. RED CINEMA digital cameras) are also "printed" to film stock to be archived long term?
Drives can be added or subtracted at any time so the system grows as your data grows and unlike traditional RAID you can mix and match drives of different types and sizes. Because it is not a real RAID implementation you can pull out any data drive in your array and read the files residing on that drive in another system. Data is protected against drive failure by single or dual parity drives and using SSD's for caching is even supported.
For anyone that is mildly technically inclined and enjoys DIY solutions, I would recommend at least checking out UNRAID before purchasing a Synology NAS (or similar).
I've been using it for a few years, and it's survived one drive failure (which I replaced with a disk twice the size). I originally had it on a Ubuntu server but have since migrated to running in openmediavault (which does all the setup for you), on a vm within proxmox server.
> Please do not use a comcast.net email address to purchase a license key.
Anyone know what the problem with comcast.net email addresses is? I know that in general it is not great to use addresses tied to your ISP for long term things, but that that applies to most ISPs so I’m guessing the issue is something specific to Comcast.
> Please do not use a comcast.net email address to purchase a license key, they blackhole our email.
Super simple and short (assuming that explanation from a sibling comment is correct). People are much more likely to oblige/agree when the reason for a rule is stated.
The community is supportive, templates are created and updated quickly, and there are several youtubers that make great, easy-to-follow videos for care and feeding your server. +1.
You can expand it but not by adding a single disk to a parity set (going from five drives to six), only by adding a set of disks (adding another set of five or a similar combination).
Because I don't need web interfaces or more attack surfaces
Yet this setup doesn't notify you of that happening, so you'll have to manually keep an eye out...
This first sentence is a big one. For a while now we have been making a yearly printed album of our favorite family photos. The first one was a gift for our parents but we liked it so much we continued the project.
When close friends and relatives visit, it is much more likely that we'll pull an album off the shelf and enjoy the memories together than we would digitally go through a year's worth of pictures.
It's a fun family project and you end up with a nice curated keepsake that will last a long time.
I also store a copy of each album sealed in a mylar bag which should keep them nicely for any curious decedents long from now.
https://forum.photostructure.com/t/front-page-of-hacker-news...
I had so many of my family, friends, and beta users ask me this question that I decided to do some research and write it up.
If you've got suggestions, or find any bit confusing, I'm all ears.
Getting close to upgrading my old 2 drive Synology at home that has been running non stop for about 7 years without fail and was just going to do another 2 drive with some new large capacity drives (Synology 720+).
If you run RAID-1 in a 2-slot NAS, and you run out of space, you've got two options:
1) buy another NAS
2) move off of RAID, and have 2 distinct volumes. You'll have twice the storage, but no spindle redundancy.
If you've got 4 or 5 slots, you can throw in 2 larger disks, and use the prior disks either for backup, or use the new disks for incremental additional storage. You've just got a ton more flexibility.
"Consider getting a NAS that has 4 or more drives to offer redundancy and support data integrity checks"
I may go for a 4 bay, but I never upgraded capacity in the last 7 years (3TB mirrored) and would probably go 14TB mirrored on the new one, so seems like a waste.
Wow, yeah, if your storage growth is that stable, and there's a big cost penalty for 4+ slots, then a 2 slot NAS would seem like a good idea.
Main thing I will use more capacity for will be more surveillance camera footage retention and more and higher res cameras.
Would that result in blurring? Perhaps the example shown in the article (where each individual bit flip resulted in major, obviously abnormal changes) was atypical?
There are almost 10²⁰ ways to flip 3 bits in 1 MB. This approach seems unfeasible if any more than 1 bit has been flipped.
Even the ability to correct single bit flips would be useful though. I can also imagine much faster ways to do it... for instance, since the tops of the images are fine, the problem is likely in the 8x8 block where the corruption began (and the user could point to this).
May be right but mdisc grade storage is way more better on the long run than a simple hdd. Accidental deletion, failures, ransomware are all much larger problems than many thinks when it comes to mediums other than the read only optical disks.
Each time that optical disks are considered, they talk only about the organic dyes of the Cd-r-s. HTL blu rays don't have issues like this, and mdisc blue rays are even more durable.
I've recently bought a synology to store my photos. But I'm questioning my reasoning for this at the moment. The upside is that I have a central place to store photos so they aren't scattered around on different laptops and external drives.
But the downside is that it's all in one place and even with backups it's more vulnerable to a single point failure.
So maybe there is some merit to storing photos on smaller external drives and having a few cheap 500gb drives for each year (and make some copies of course). And then maybe pick the best photos and burn to some HTL BR discs. The biggest problem with external drives is incompatible filesystems and fat32/exfat being susceptible to corruption.
Because honestly I'm not going to take another photo in the year 2020, am I?
So instead of diligently backing up my NAS, securing it from hackers and viruses, hoping btrfs devs wrote unit tests etc (1), I can make a few redundant copies on different hard drives and usb sticks and leave some at my parents house etc.
Theoria Apophasis, a youtuber ("the angry photographer"), has a few rants about long term data storage and archiving. He is a bit eccentric but imho makes some good points. Btw, don't watch if you don't want to have nightmares about your hard disk dying any minute now, ignorance is bliss.
- Methodology to protect your data. Backups vs. Archives. https://discussions.apple.com/docs/DOC-6031
- DO NOT LOSE YOUR DATA! Please be wise about DATA STORAGE https://youtu.be/o99hwegrvJQ
- Data-God Photographer: Part 1: Backups, Archives & Redundancies, Hard Drives & Optical https://youtu.be/scLMP9gm--M
- Angry Photographer: Video 1. HOW Hard Drives WORK, why you MUST fear & hate them https://youtu.be/uKGsNoUZAO8
- The ONLY SOURCE for long-term Archival DATA PROTECTION is... https://youtu.be/qbxaPc2Xf5M
(1) Tbh, I also have no faith in synology software, my user experience so far has been pretty disheartening. Photo station & moments are horrible, the dir is hard coded, you can't choose where your photos are stored which is just ridiculous. The backup software doesn't have an archive option for deleted files (like the CCC "Saftey Net" or rclone backup), it can only delete or keep the files. Also, for example the AFP protocol loses the file modified date: https://discussions.apple.com/thread/7547857
My understanding is that they're putting a shiny wrapper on some open source software and reselling for a premium. Basically this https://xkcd.com/2347/ I want to be that happy enthusiastic person in their promotional material, but the experience is not reassuring. And supposedly everything else on the market is even worse (plenty of Drobo horror stories online).
Agree, they are unusable. I still don't understand the different user permission settings for photos and everything else. Phtos are somehow separate.
That said, I just don't use it and do all of that on external computers, using photo software.
https://community.synology.com/enu/forum/11/post/122691
And frankly this type of shenanigans makes me seriously doubt the competence of their software team.
The filesystem cannot truly know if a file being modified was intended or accidental. Sometimes a software error can wipe or corrupt a file, and a filesystem can't detect that, even if it can detect bitflips.
The data architecture I use combines, in order of importance: (1) borg, for incremental deduplicated backups onsite and offsite; (2) syncthing, for duplication and synchronization across devices; (3) fim, for managing file integrity; (4) rsync, local copy to another drive for convenient restores; and (5) git, in places where commit history is actually useful, like code or dotfiles.
fim combined with incremental backup solves this problem, is filesystem agnostic and painless to migrate around, and is for some reason completely unknown and obscure: https://evrignaud.github.io/fim/
What you could do is put your photos subdirectory inside another directory, something like below.
data/
├── .fim
│ ├── settings.json
│ └── states
│ └── state_1.json.gz
├── photos
└── recordsThat looks overly complicated
Why is there no single tool that can do all of those?
I think one tool could be much more efficient. Especially since they could reuse the databases. With one database of hashes of each file, it could find the changed files and then incrementally only backup those. No need to have one program to search changed files to copy them, and then again have one program to search changed files to hash them
I mostly use rsync. Occasionally I destroy my backups by calling it wrongly, like missing a trailing slash. That could not happen if the copying and file integrity checking was combined in one tool
>fim combined with incremental backup solves this problem, is filesystem agnostic and painless to migrate around, and is for some reason completely unknown and obscure: https://evrignaud.github.io/fim/
That looks very useful
But it stores the database as json? That is a bad format to store file names. And written in Java it is probably going to be slow (I have over a million files in ~, and multiple copies of it in the backups, so I care a lot about performance)
So many emojis in the commit log. Is that really necessary nowadays?
In practice, it's simple enough. I have a backup script that does all the heavy lifting for borg and rsync. Syncthing operates independently in the background, and I rarely log changes with git. The most tedious part of the process is logging new/changed/deleted files with fim.
> Why is there no single tool that can do all of those?
Good question, and I totally agree. I was actually looking for a solution that could do deduplicated incremental backups (versioning, like time machine) plus file integrity management. I couldn't find one. Nobody seems to care much about file integrity, and nowadays I'm actually unsure if this is a real problem for anyone except paranoid geeks.
Could also use a good frontend GUI that could notify you about changed files, allow you to easily mark directories for different levels of "tracking", easily browse past versions and restore them, etc. The CLI isn't doing any favors here.
> I mostly use rsync. Occasionally I destroy my backups by calling it wrongly, like missing a trailing slash. That could not happen if the copying and file integrity checking was combined in one tool
Try borg-backup, it'll be a massive improvement over plain rsync when it comes to backups.
> But it stores the database as json? That is a bad format to store file names. And written in Java it is probably going to be slow (I have over a million files in ~, and multiple copies of it in the backups, so I care a lot about performance)
Gzipped json on-disk, but I haven't noticed any problems with it over the last few years. I doubt java is bottlenecking the performance at all: it's limited by IO and CPU throughput for hashing the files. For everyday use, you don't need to hash every file in entirety to detect changes. There's a "fast" mode that checks a couple of blocks. You can also operate in a subdirectory of the repo to ignore files outside that subdirectory.
I also don't just use one massive fim repository. I have a .fim/ for each directory I care about (types of data) that a secondary scripts loops through for checking the status. That weakens the integrity guarantees since you can't log the movement of files across directories, but makes it easier to reorganize things.
> So many emojis in the commit log. Is that really necessary nowadays?
Yah, that's the most emojis I've ever seen in a git repo. Maybe 2017 was a different era.
Of course this is only as good as the interval you use for your snapshots. Fim seems like it would be a lot of extra work to do for every file/folder, but I could potentially see using Fim though for important files (e.g. tax records?, password vault?)
Files that change frequently (password database) or won't totally bork out if there's a bit flip (text notes and org files) don't really need their integrity managed manually.
For "current year" photos i store them on my NAS, backed up nightly, locally and remote.
Every year i then archive the previous years photos onto an identical pair of M-disc BDXL discs (100GB), and store one copy locally and one copy at a "temperature controlled" remote place. I use no compression, encryption or archiving. If there is degration of the physical media i want to minimize the damage.
Along with the archive BDXL media, i store an external drive which contains a complete copy of all photos. I also update this drive every year, and run a non destructive badblocks test on it as well as a long smart test before updating it. The drive is then updated, and rotated with the remote one (typically when storing new M-disc media), and i repeat the process for the retrieved drive.
Again, plain ext4 filesystem (i HOPE FAT32 will finally be dead in a couple of decades), no compression, archiving or encryption.
I never delete the photos from my NAS. The archive is simply "disaster recovery". I only archive photos. Any document will probably have little value in a decade or two.
I'm also fully prepared to migrate my entire archive onto whatever the "next big thing" in archiving becomes. No archives are truly forever. Optical media can go away, USB 2/3 will certainly become obsolete at some point. There's no point in having an archive i cannot access.
And yes, i have considered just creating physical copies of the images and storing them in identical photo albums, but sadly nobody "develops film" anymore, and everything is printed, which also degrades over time, and unlike film of old there are no negatives to reprint images from.
Btw, do you think there is a difference between 25GB single layer and 100GB M-Discs regarding the media robustness and safety of data?
As i understand it, the 100GB discs add another data layer, and assuming it works like all other optical media, this is achieved by angling the laser at a different angle, so logically a scratch on the physical media will incur read errors on another layer.
Apart from physical damage, the media density is higher.
I would assume that the 25GB discs have a longer lifetime than the 100GB ones, but I have no illusions that either of them will last a millennia. If the media is still readable 1000 years from now, let’s hope they still make optical drives to read them :-)
I’ll be happy if they last 10-20 years, and considering some of the early CD ROM disc that were burned are still readable despite using dye (which degrades) I’d say there’s a good chance of that either size will survive just fine.
I store mine in jewel cases, in a dark closet at room temperature. The worst enemies are temperature, light and humidity.
I have recently started using BTRFS snapshots on my playground server, and my god is this awesome. I get ZFS is quite special, but why is BTRFS not default FS?
Like, if try to cause bitrot manually, and see how each recover?
There's this: https://askubuntu.com/questions/406463/how-can-i-flip-a-sing...
Incidently: I thought ZFS was ancient, and btrfs was a fairly new project, but I was wrong:
- ZFS was available in 2004: https://en.wikipedia.org/wiki/ZFS#Sun_Microsystems_(to_2010)
- btrfs was available in mainline linux in 2009: https://en.wikipedia.org/wiki/Btrfs#History
> why is BTRFS not default FS
Perhaps because there are still some sharp edges around btrfs, particularly around RAID56 support. See https://btrfs.wiki.kernel.org/index.php/Status for details.
Good thing I had syncthing running.
Also, maybe Synology will manage its drives better than I did in my multiple-purpose desktop machine.
Like, I've been using ZFS for about 10 years without a problem -- I had a disk in a mirrored setup go bad that I didn't notice for awhile, and when I replaced and rebuilt, I had a couple of files which had suffered corruption.
But that's just my personal anecdote; I'm quite confident there are people using BTFS who can quote similar track records, and I assume there are people who have seen ZFS shit the bed for one reason or another.
In contrast, I never had a single non-hardware-related problem with ZFS in maybe decade of using it for NAS. It's totally rock solid! I've recently figured out how to do unattended root on ZFS setup to provision my PCs and never going back to btrfs no more.
The file systems couldn't be restored with the official suite of tools. It was a backup anyway and we restored the data from secondary backup.
I'm not positive that there is any other file system that I've ever lost significant data on.
https://cloud.google.com/storage/archival/
There is also Amazon Glacier
https://aws.amazon.com/glacier/
They market these services towards SMEs, but one-man-shows/solopreneurs can use them too.
In every case those became a problem over time either due to degradation of the media or simply because the underlying tech became more rare or unavailable.
So about ten years ago I flipped my approach. My long term storage is my current file server and for the backups I have no expectation of them lasting more than a year or two but that's fine.
The file server is up and running 24x7 so it is not bit-rotting in some drawer. It runs ZFS so integrity is guaranteed. ECC RAM. 4-way mirror for every pool so there's plenty of redundancy. The only drawback is having to stock it with enough storage to hold everything, but that's the tradeoff. Storage gets cheaper while backup media doesn't get particularly more reliable so it's a good tradeoff.
I still do backups of course (while ZFS and ECC guarantee integrity, nothing guarantees I won't accidentaly rm a file) with zfs snapshot and send, but these are only for disaster recovery not archival so I don't need to lose sleep worrying how long they last.
Pick an off site backup system where you have enough space to store every version of every file, with infinite retention.
What you want to avoid is the single most common error: you accidentally delete or corrupt something, then several years later you notice.
This might sound unlikely but it is I assume more likely than theft, fire, and other reasons we keep our backups off site.
Whether "retain all copies of everything, forever" is actually viable can vary with budget and the size of ones library of course. But storage is pretty cheap these days.
Instead, I just set the retention policy to a year. Since most of my backups are photos and archival stuff like that, I rarely delete those without confirmation, so it's not a problem.
For other things like documents, I have Nextcloud, which does keep the last N versions, so it guards against the scenario you describe, true.
That said, I too have a few TB of online storage, but my problem is the reverse: I will sometimes capture a few GB of photos, import them for safety, the nightly backup will run and back them all up, and the next day I'll delete most of them as rejects. Those rejects, I don't really care to keep, so I wouldn't want blurry or otherwise useless photos taking up space for ever.
Support for borg remotes was recently added, which I think will be very useful.
I try to keep it very simple. I use cron to rsync --backup to another computer and a NAS.
I do not use Raid. I want to be able to take any disk and mount it on any computer, possible using a USB adaptor.
There is no need for instant synchronization. When I upload pictures to the photo album, I wait a couple of days before deleting them from the camera.
I set up a screensaver on a desktop to display random pictures from the NAS backup. We like to see the pictures that way but it also have the advantage that we notice when the NAS is not working and I do check that I see a recent picture once in a while.
I now make certain that I use photo album software that organize pictures using directories and filenames. I once had a crash and recovered the actual photos but not the database with tags, names etc. So I had to name all pictures again.
That can be problem if you want to use different operating systems and don't want to use fat32/exfat.
It seems like a good enough solution for me to not worry or think too hard about it. Syncthing isn't the most user friendly software though. When I upgraded from 2TB to 8TB drives, I accidentally a setting and ended up with about 1TB of duplicated "sync conflict" files. If I wasn't a salty software engineer comfortable writing arbitrarily complex scripts with free reign to delete files, it would have taken more than a weekend afternoon to clean that up...
I used various arm single board computers previously, with the hard drive connected via USB. When I started using syncthing, though, it was literally taking days to hash through all the files, only going 1-2 MB/sec. The UI wouldn't work until hours after restarting the daemon because it took so long just to build its data structures in RAM. The baytrail CPU has enough oomph to tear through the hashing, and it's just nice to have expandable RAM and SATA.
The reason being is you can decide how long to keep older copies of files, just in case something changes on a file, but you don't happen to realize the incorrect update for a couple of months and want to go back to an earlier version.
To give you an idea of what is possible, here is the snapshot configuration I use: Keep no snapshots older than 450 days Keep 1 snapshot every 90 day(s) if older than 180 day(s) Keep 1 snapshot every 30 day(s) if older than 30 day(s) Keep 1 snapshot every 1 day(s) if older than 1 day(s)
The 3rd hard disk is a cold backup in case syncthing has a common mode failure that takes out my redundant drives. I want its filesystem to be as boring as possible. For me that currently means ext4. I don't want to have to manage and worry about another tool that optimizes away data transfers, when there's value in the data transfer and plenty of copy bandwidth going SATA to SATA.
For example, say you are using a financial application like GnuCash on your desktop and it has bug in the software that gets triggered causing a bad write to the file. (power goes out, OOM, whatever..) syncthing will happily propagate that changed file that contains the bad write to your server and when you copy to the 3rd disk you will also propagate that changed file. You deleted what was on the third disk, so now you no longer have a "good copy" of the file.
If you add some sort of snapshotting into your backup routine, then you would be protected from this, because even though you would still propagate the change to the most recent backups, you could still go back to an older snapshot (maybe a week or a month ago or whatever) and pull back a working version of the file.
Also, I have enough disks laying around that I usually have a 2nd cold backup to wipe. So, I have some inadvertent temporal redundancy going back a year or two.
When I make a conscious decision to copy files over to my storage, I have piece of mind that I'll get physical redundancy within hours, and will get picked up by my cold backup eventually.
You can't. But you can make multiple backups, each one reduces the chance of losing your files. Store at least one backup off site.
Also, media dies over time. Buy new hard disks annually and make them the new backups.
The old drive itself failed within a year.
So I'd feel pretty safe copying one drive to another every 5 years.
When it comes to digital images (I don't have enough to bother), archive-grade printing to paper seems like the obvious first-choice option -- IF it's done properly.
Then, it would cost you $0.44 to download all 45 GB on the same day.
For $25.64 total (over 12 years), they store your data with significantly more redundancy than you when you put one copy of your dataset on one hard drive.
I chose B2 in this example because they're cheaper than blob storage from the main clouds, you're unlikely to get banned for an unrelated reason (cf. Google), and their pricing model is simple to understand.
Assuming ~120 MiB per compressed CD album, and ignoring the futzing about SI and Binary units, you can store at least 8 CDs worth of ~192 kbps music (~1 GB) in B2 for $0.005/month, and since the first 10 GB/month is free, your first 80 albums are stored for free. Then, each additional group of 8 albums is another $0.06/year.
If you're still unconvinced, and prefer the particular characteristics of control, convenience, and no direct monetary opex costs that personal self-managed storage affords, then consider that for an extra $25.64 over 12 years ($2.10 for storage-at-rest, yearly), you can have another copy of your 45 GB in the cloud, which significantly reduces the likelihood that your dataset is damaged.
You can even think of it as insurance, but with the extremely desirable property that you get your actual data back, and not just some other kind of compensation.
I currently back up to two drives at ingest, one being a per-year drive (e.g., 2020 jobs source media) that is rarely accessed. For major projects, I factor in one or two 1-2TB drives that receive an extra copy of source media and then are stashed. Not perfect, but so far so good.
I know these features are provided by some newer file systems like zfs and btrfs, but those are not used by most consumers and are mostly used on servers.
OEMs could choose to solve this for operating systems and maybe they do on internal storage, but it doesn't help with external storage, which is once again what 9 out of 10 people consider to be a "hard drive".
Hard drive manufactures therefore have a strong incentive to solve durability as best they can inside the drive units themselves. I don't actually worry about bitrot too much because I assume there's some sort of propietary error correction mechanism at the firmware level. I expect my data is either going to read out flawless, have gaping holes in it, or have the whole disk unit fail. I'm not going to blindly trust that such a mechanism exists, but it weighs into my risk analysis when making decisions on how to best protect my personal data.
Don't treat retention as free. Its either your labour and disks or reliable files for an SLA but its $ no matter what.
Be philosophical when dateloss happens
https://arstechnica.com/information-technology/2014/01/bitro...
How exactly does a NAS protect data from bit rot? Data scrubbing?
I assume a nice GUI to tune parameters?
The GUI _is_ pretty slick.
Even if Google had reasonable customer support, relying only on a cloud backup is unwise.
UX being what it is, my aunt just lost most of her iCloud photos because she thought she was cleaning up her laptop to donate, but iCloud synchronized all the deletes up to her account.
My father "clicked the wrong button" and lost his entire Picasa library a while back.
More copies are better.
https://www.imyfone.com/ios-data-recovery/how-to-recover-pho...
Pricing plans and other terms of service can (and will) abruptly change.
There is no guarantee that Google will even exist 20 years from now.
If we are talking about data that is to be handed down for generations, I certainly wouldn't trust Google as a solution. I do agree though that setting up self-hosted local/offsite backup, the kind described in the article, is out of reach for most users. Maybe redundant copies on every cloud storage provider will do the trick?