Backing up 18 years in 8 hours
chromakode.com
chromakode.com
Files since 1999 are all still in a folder on my local drive. I've made an effort not to lose the photos and chat logs. My iTunes database dates back to 2002, and I'm disappointed that I lost my old SoundJam MP database that I used for 2 years before that.
Unfortunately, my archives continuously come under threat from forced upgrades. Upgrading iPhoto to the Photos app will lose a lot of metadata. The latest versions of iTunes removed USB sync for contacts & calendars, which really bothers me when I want to maintain my files without the cloud.
Many old documents are now unreadable (e.g. Clarisworks). I've realised that simplicity is essential for keeping long-term records. Keeping a copy as plain text is important for preservation. Just because "there's an app for that" today doesn't mean that 10 years from now, the app will still work. This is (especially) true for companies that should know better, such as budget tracking tools.
I recently went through my old records to compile a list of every rock concert I've been to, along with the location and price when possible. With heavy use of archive.org, my iCal, and my iPhoto library, I figured out all the dates, but it was a major effort. Most of them were in the last 10 years.
My dad still has his PhD thesis on a magnetic tape. Anybody who knows how to read that onto a PDP-8 should get in touch!
It's really disappointing when software uses unnecessarily complex file formats. Chat logs are an important example: I was dismayed to discover that my Yahoo Messenger chat logs on Windows are stored XOR-obfuscated! [1]
There's a lot of compromises to be made between the convenience of making backups and their long term viability. My hope is that by greedily saving all of the bits, there will be enough context to make sense of the data if I really needed to, even if it is an intensive process.
A good alternate strategy would be to scoop up the most likely interesting files when processing the backups and re-encode them in more future-ready formats, as derefr suggested.
Batch re-encoding those files at the cusp of losing easy access to them turned out to be a lot of possibly error prone work. So to future-proof important documents, now I do that re-encoding continuously up front, whenever those files are saved.
But for long term viability, on top of backups, the files themselves have to be in a format usable by programs that will be around for a long time on platforms that won't disappear.
Practically, that means using simple file formats (TXT). Or else using a program (LibreOffice) that creates files in an open format that can be easily re-encoded up front into multiple formats like DOC and PDF. The MultiFormatSave extension for LibreOffice makes it easy to save into multiple formats for that purpose.
After all that work though, I just found out LibreOffice can open ClarisWorks files. lol...
This is exactly why some people are so vocal about open formats.
Regarding your fathers PhD thesis, take a look at http://www.pdp8.net/ I'd be surprised if they cannot help or point you to some people that can.
Regarding the article - I'm not sure I understand the concept behind storing a set of the recovered files in a Cloud service. Surely a second HDD with the same contents stored at a second location would be cheaper in the long run. The HDD would last 20+ years if stored correctly - you can't guarantee that your consumer level Cloud service will still be around even in 10 years time.
I wouldn't trust a HDD for 20 years. Keep a backup mirror locally and a backup offsite (cloud or at a relative), what matter is that the transfer is automatically running in a periodic manner.
I hazily remember having the Outlook PST files as the mail archives and then later not being able to recover most of the metadata of the mails, not even programmatically, as the PST was designed even before the "internet standards" were something they'd worry.
GPG also removed the support for the old encrypted files and the old keys since 2.1: https://www.gnupg.org/faq/whats-new-in-2.1.html#nopgp2 which signifies the dangers of using it in archival contexts. The disk encryption tools can be even more problematic, being dependent on the particular OS versions.
I have most of the stuff since, though floppy backups were erratic. I copied the 5.25" floppies to CD, and then to hard drives, before discarding the old computers.
I try to save most files as jpg, pdf, mp3, mp4, or plain text, figuring those are the most future proof.
(They basically introduced pdf, because postscript was too open and other companies started beating them at it.)
https://en.wikipedia.org/wiki/PDF/A
http://www.preforma-project.eu/media-type-and-standards.html
The LOC has good guidelines for archiving all sorts of formats.
”Upgrading iPhoto to the Photos app will lose a lot of metadata.”
I’m still on Mavericks but will upgrade at some point. This makes me nervous though. Would you care to elaborate on this issue?
https://support.apple.com/en-us/HT204478
One thing to bear in mind is Photos will create a new photo library, leaving the old one untouched. The metadata will still be there, but good luck getting it out.
There is/will be a niche market for folks writing software to convert legacy formats into open/current formats.
Remember how all of a sudden shops offered to convert your VHS-tapes to DVD? I bet they were making a killing.
On the one hand, I can imagine "cryonically" preserving a disk image for a later, better filesystem recovery program to come around. (This "cryonic" approach would give even better results by preserving bit-level analogue flux recordings of the disk platters, rather than relying on the output of digital reads from the disk heads.)
On the other hand, the longer you leave these disk images as dead blobs of data, the more layers of legacy container+encoding formats you'll have to try to get your system emulating when you finally do want to pull the files off. One day your OS won't have drivers for reading e.g. FAT16, or zfs2, or ReiserFS2.11, nor will it be able to parse out the meaning of an MBR-partitioned disk. Reaching back through Linux kernel archives for something old enough to understand those things, will result in a kernel that won't boot your PC. You'll end up having to do something convoluted with qemu just to get your disk read.
Personally, I'd much rather throw out all the intermediate containers I can, as soon as I can: not just extracting files from the disk's filesystem, but further extracting files from any proprietary archive formats on the disk (using the extractor tools probably installed on the same disk), and even canonicalizing containers like AVI by remuxing them into modern extensible formats like MKV. The goal being to give a file-level deduplication process the best possible inputs to work with, most likely to match: not just for space-saving, but because reducing "junk duplicates" helps greatly in actually finding anything in all that mess, let alone organizing it.
> One day your OS won't have drivers for reading e.g. FAT16, or zfs2, or ReiserFS2.11, nor will it be able to parse out the meaning of an MBR-partitioned disk.
Maybe in 10-20 years.
Perhaps some things are best left in the past.
If your backup plan involves mounting a disk image on a VM, I strongly advise you to test it before you need it :)
Creating new Linux VMs expressly for the purpose of reading particular file formats is a much safer bet. It is unlikely that the ability to conveniently emulate a basic i386 system with block storage will go away in the next few decades. My assumption is that any formats I can read on Linux today will be readable in the future, as long as I have a copy of the source / binaries. This is why I included a copy of lrzip's git tree with my backups -- everything else is in a standard Ubuntu install.
That being said, a good guess would be that the most interesting data will be media files (especially old photos) and documents. For that data, your advice of collecting and re-encoding the files is wise. For the purpose of discovering the media files in these backups, I found my favored approach to be a brute force recursive search for file types. Exploring the original structure of the filesystems was interesting, but my intuition for where the valuable data was usually proved wrong.
Of course, things are never as simple as they should be. LaTeX is a markup language that through macro expansion ends up expanding into TeX. Since TeX 3.0 in 1989, Knuth has attempted to keep the TeX system stable. Since then TeX documents should produce the same output, pixel for pixel, as they do now running on the current version 3.14159265 (yes, the version number is converging to pi, there won't ever be a TeX 4.0).
Few people, however, produce documents in plain TeX--the LaTeX markup is so much more convenient than the lower level TeX. LaTeX has been slowly evolving and there are some backward compatibility issues, but they are minor. The first release of LaTeX seeing general use was described by Leslie Lamport (the creator of LaTeX) in his 1985 book[1]. That version of LaTeX, 2.09, can still be processed by today's LaTeX 2e in compatibility mode. LaTeX 3 is supposed to supersede LaTeX 2e someday, but it's not clear how many more years that will be.
Since LaTeX is open source the distributions from TUG (the TeX User's Group) are easily obtained and they have all the historical versions of LaTeX and TeX available.
This all makes LaTeX/TeX seem like one of the best ways to maintain a document's source for the future. A few tips:
- Fonts can be a problem because fonts evolve. Either use something like TeX's extensive collection of "built in" Computer Modern fonts or save the font files along with the source of anything that you might want to work on in twenty years.
- LaTeX has a wide number of very sophisticated third party extensions. Along with the document source it would be a good idea to keep the contemporaneous versions the extensions used by the document. (These extensions are just files of additional macros.)
- If one is only interested in the typeset results of a LaTeX document, use a LaTeX extension like pdfx (or Adobe Acrobat) to generate a PDF/A version of the document (LaTeX programs normally generate pdfs). PDF/A is a pdf specification from Adobe that is intended for archival use and rendering far in the future (fonts are embedded in the output, etc.).
- Pandoc can convert LaTeX to a wide number of alternative formats (HTLM etc.) with some success depending on the target language and the document complexity.
[1] http://www.amazon.com/Latex-Document-Preparation-System-User...
PDF/A is a pdf specification from Adobe that is intended
for archival use and rendering far in the future (fonts are
embedded in the output, etc.).
Maybe they state that, however as long as there is no reader that is working on the newest hardware this is just wrong.And the Specification for PDF/A-1 - PDF/A-3 (a,b,u) is really really long and hard to implement correctly. I doubt that 70% of the solutions (even for money) incorrectly parses the spec. Actually I doubt that only Adobe Acrobat (Reader) would actually print a 100% correct PDF/A-3a file to the screen.
Actually this is the (inofficial/offical) spec:
http://www.adobe.com/content/dam/Adobe/en/devnet/acrobat/pdfs/PDF32000_2008.pdf
It should be really close to the iso one (http://www.iso.org/iso/home/store/catalogue_tc/catalogue_det...)Your idea of a VM container is interesting--I hadn't thought of that, but what about the software to run the VM. Does VirtualBox stay backward compatible over the next 10 or 20 years? Maybe the way to go is Markdown, that is so easy to read even without processing it.
Re: markdown or similar like restructuredtext, I don't have any experience with advanced markup in them, but I think you'll have the same plugin problem as with LaTeX. The readability of the source isn't that important, since you'll probably have the rendered output as well, and that is,or should be, copy/pastable if needed.
Now, if you are a savvy programmer and you archive the (open) OGG spec, you could argue that you can write your own custom OGG -> MP2073 encoder at any time to recover your ancient music. But that's not most of us.
Another example that comes to mind is the many MS Office replacements I've used over the years. I have a small trove of old documents in old open formats used by StarOffice, OpenOffice, & others that are a pain in the butt to access and don't always render correctly, while my ancient .doc files from twenty years ago still open in two clicks.
Correct. MP4 is not an open format.
From [1]:
> MPEG-4 contains patented technologies, the use of which requires licensing in countries that acknowledge software algorithm patents. Over two dozen companies claim to have patents covering MPEG-4. MPEG LA licenses patents required for MPEG-4 Part 2 Visual from a wide range of companies (audio is licensed separately) and lists all of its licensors and licensees on the site. New licenses for MPEG-4 System patents are under development and no new licenses are being offered while holders of its old MPEG-4 Systems license are still covered under the terms of that license for the patents listed (MPEG LA – Patent List).
> AT&T is trying to sue companies such as Apple Inc. over alleged MPEG-4 patent infringement. The terms of Apple's QuickTime 7 license for users describes in paragraph 14 the terms under Apple's existing MPEG-4 System Patent Portfolio license from MPEG LA.
Cheap Internet access (i.e. a local dial-up number) didn't come to the area where I grew up (a small, midwestern town) until about 1996. From 1990-6, I was heavily involved in the BBS "scene". Being a teenager with lots of time but little money, I ran a BBS but couldn't afford to pay for all those door games. Instead, I learned to reverse engineer them and write "cracks" (binary patches, basically) to "unlock" those games. In the middle of the night, I'd dial into far away BBSes and upload them everywhere that I could. I'd love to have that source code to look at again today.
I did find some of my old paper notebooks in my mother's garage a while back. These are the ones that I took to school with me. Instead of doing schoolwork, however, I'd spend my time in class writing code (on paper, by hand) and then I'd type it all in to the computer when I got home after school. It was very neat to find those and look back at code I wrote ~25 years ago.
For the last few years, now that we all have a camera in our pocket at all times, I have kept backups of every picture I've taken but there are probably thousands of photos that I took before then that I'll never see again.
With storage as cheap as it is, I've resolved to never again "lose" any of the photos or videos I've taken. It'll be amazing to look back at them in another ~25 years or so.
Happend to me. Encrypted an archive, and the funny thing is I did save the password into a password manager, but I didn't label it descriptively enough, so when I tried decrypting the archive I couldn't locate the right password. I could have just tried them all, but I had a copy of the files from the archive in another place so I just deleted the encrypted archive.
To clarify: I didn’t primarily choose my backup application because it encrypts the backups but because it seemed like a reliable application with a solid user interface. (In Arq, encryption is not optional.)
Joking aside, I do use some things like hashes of works of literature and other important things to me as keys and archive hashes as their own passwords. I also send emails to myself explaining my thinking when I make certain decisions. It's like commenting my life.
And if I lost my memories, would I really selectively miss the data that belonged to former me? That's a scenario to think about.
Not surprisingly, when I visited Moscow 15 years later, my old hardware was gone, and not a shred of knowledge remained of where it might had gone. My relatives conceded that it might have been stolen. I hope the thieving computer obsolescence club or history museum will enjoy my pieces.
If we are so concerned about our privacy or encryption, we should not make such mistakes.
Great job
(To make a point you don't have to show full coordinates)
In the future, it's really important to attempt such notices privately first. Even if the information is out there, making that knowledge public in a popular spot greatly increases the exposure. I'm sure that with your skillset you could have a direct line to contact me in less than a minute.
My point is. The technical aspect of moving data around is actually not that challenging compared to the "business logic" of organizing your old files.
I find this an excellent metaphor for programming. In particular: low-level debugging of high-level code is inevitable, so choose your abstractions carefully.
That said, this is the one sentence my 1st grade teacher always kept repeating and stuck with me ever since. (it was, at the time, about pens falling from desks)
Bank A shouldn't be in the same region/area as bank B in case a desaster strikes and floods the vault(s).
Then play that scenario through whenever you update your keys/data.
My macbook disk crashed, and my timemachine wasn't working for some time.. It was 200GB of data.
What I did was connect the disk to a linux machine, and use ddrescue for a month 24/7 to try and read the bad disk.
I lost 0 bytes.
[1] that thing had pseudo modules, a visual tracer, live indent and lint (limited but still).
It's also so strange when co-workers talk about how their kids are not just playing games but also creating mods for their games. When I was kid I just drew pictures, played with lego and played video games.
Edit: updated the post to mention this as well.
From https://cloud.google.com/storage/docs/nearline:
> Data is available in seconds, not hours or days. With ~3 second response times and 1 cent per GB/month pricing
This is in contrast to 3 hours of Amazon Glacier [1].
https://twitter.com/apaprocki/status/550432891201941504
In case anyone else has a stack of long forgotten various QIC formats -- don't lose hope, have patience and restore them :)
I believe the way to get the best reliability is to keep multiple replicas of the archive (git-annex's great here) over storage options from as many different vendors as possible, and hope the chance of everyone going down at nearly the same time is too low. And probably schedule some periodic fully-automated checks to minimize the chance some service had degraded and had silently lost the data.
And as technologies evolve, in 5y there will probably be cheap 30TB HDD, so it will make sense to consolidated, etc
I know this is clunky, but I'd print out the key (or possible laser-cut the key!) onto durable media and store it in a safe / safety deposit box etc.
It's going to be annoying to manually type it in, but it's a reliable method of last-resort.
Laser-cut the key represented as a QR code, paired with classic textual hexadecimal (not even base64, to not have to worry about "o"s and zeroes) representation. Problem solved.
Be careful. There's a tendency that if one drive fails, the other will follow very soon. Especially if the drives are from the same batch (I was bitten by this), but not necessarily. I'm not expert on the topic, but I believe that when planning for long-term HDD reliability, buying drives from multiple different vendors is a must.
I also like that you encrypted your backups, although in doing so you probably made Senator Dianne Feinstein cry.