Archivists are uploading hundreds of random VHS tapes to the internet
vice.com
vice.com
Young people probably would not know that watching VHS tapes anything could happen because they got reused.
So you'd be watching a movie and halfway through it would suddenly switch over to a shuttle launch or a music video or a documentary or something cause someone decided to record something else at that point.
I once recorded a three hour British detective show. I watched it for three hours and it got to the final scene to reveal whodunnit and the tape ran out.
It's great that they are archiving the old content but I don't miss VHS in the least.
I ran into one of these on Youtube just a couple of weeks ago. The reaction in the comments was something less than amused.
(Bittersweet memories of taping over some of Mom's seemingly non-important video in order to record...what was it, the first X-files episode? A _Wings_ episode about a favorite jet? Something like that. But man, she was not happy.)
#!/usr/bin/python
from random import choice
c=["Miss Scarlet","Mr. Green","Colonel Mustard","Professor Plum","Mrs. Peacock","Mrs. White"]
w=["candlestick","dagger","lead pipe","revolver","rope","wrench"]
r=["kitchen","ballroom","conservatory","dining room","cellar","billiard room","library","lounge","hall","study"]
print("%s did it with the %s in the %s."%(choice(c),choice(w),choice(r)))I felt this, hard. So many things I recorded ended up like that, or worse, turned into my mom's soap operas - Santa Barbara, in specific. Ugh.
My oldest files are from 1977 - proof:
https://github.com/DigitalMars/Empire-for-PDP-10
My files have gone from magtape to 8" floppy to various 5.25" floppy to 3.5 floppy to zip drives to cdroms to dvdroms, then to hard disks of ever-increasing size.
(My old hard drives are completely unreadable now.)
I'm sorry I never kept my punch card decks. I'm sure there was nothing but crap on them, but it would be fun to see what kind of crap it was.
It's just life. A lot of what may seem important with our backups really isn't. Whatever we have produced today will return to the earth, and will be reinvented by generations in the future.
It's their lives to live, not ours.
Imagine having to take care of the backup of your ancestors data on 10 generations, that'd be ridiculous.
You also don't really know what future people might find interesting. Archaeologists like to sift through ancient trash dumps :-)
I had an idea about this but realized it would be far more useful for more sinister purposes and stopped.
I want to sit down and figure it out one day given how prominent it was (and still is for many critical systems).
There was an earlier Empire I wrote in BASIC, but alas it was lost. Some of its vestiges remain in the Fortran version, such as a variable named Z6. (Variables in that BASIC were one letter and one digit.)
Even communication of a simple warning message for yucca mountain and WIPP proved how hard it was going to be to communicate danger over 10000 years from now. https://en.wikipedia.org/wiki/Long-time_nuclear_waste_warnin...
The big issue isn't the technology, it's the vast amounts of data that are being created at this point. Storage is cheap, but the labor that goes into managing the longevity of datasets isn't: it's essentially continually keeping your infrastructure up-to-date whilst also ensuring the integrity and readability of the datasets as was intended when they were first created. It implies regular checks of bit integrity, readability of your data, checking that you can restore your data, ensuring that you can access the data, making sure that you can find the data and everything is catalogued, ensuring that you have the rights and license to use the data,...
When it comes to physical archives of the past, you have to be aware of your own survivorship bias. We only have an idea of what is preserved to the extent that documents are archived, recorded and thus discoverable.
What we do not know is how much knowledge and information was lost to the past. When you look at documents, you're always limited to what's there. And when you hit the boundaries of what's there, then you may have indications that there was far more in the past, but you have to conclude: sadly that's lost. Either because it is physically lost, or because it might be somewhere in the archive but it's not registered yet in a catalogue and therefor not accessible.
That's why I think that making backups with "barely a thought" is only as effective as to the extent to which you have organized your data, used accessible / readable data formats and filesystems.
For instance, most people these days generate endless streams of photos with their digital devices, which then get automagically uploaded to cloud services. And that's great. The downside of that is that your ability to find a specific picture from 5 years ago is entirely restricted to the extent that you were able to organize and add specific metadata to that picture. Let alone, if you did take the opportunity to do so.
That's why I advise people to sit down, and take time to go through their digital albums to pick the nicest or most important pictures they have, print them out on quality photo paper in several copies and store them with labels in albums at different physical locations.
When it comes to longevity, your physical albums will still be accessible to your descendants some 70 or 100 years down the line. Something that isn't remotely guaranteed by cloud solutions.
And that's just photos. Consider e-mail or the countless of closed messaging apps you have been using these past years. And then scale the problem beyond the personal but to entirety of large organizations, many of which are required by law to keep an archive of their documents, correspondence and so on, not just for decades but sometimes also for perpetuity.
I disagree with the premise that we should spend time manually organizing and tagging our pictures all that much.
The metadata that the phone adds to pictures – time stamp and GPS coordinates – is already sufficient in a lot of cases for finding pictures that I look for.
And where that metadata is insufficient, improved search powered by machine learning will come to the rescue. And not just tomorrow but even today.
Just the other day, a few weeks back, I was standing in the kitchen that I share with two other people and I wondered to myself whether the kitchen knife in the dishwasher was mine (I’d bought a new one a few days prior but couldn’t remember what it looked like). I take a lot of picture of random stuff and mundane things, most of which I never bother to organize or tag or anything. I pull up my phone, search my photo library for “knife” and lo and behold, I did take a picture of it when I bought it and my phone has recognized the object in the photo to be a knife so it was able to find it for me.
Important files and photos I do organize. Specifically for three reasons:
1. Ease of access.
2. Grouping related data together.
3. Tying photos and other data to abstract concepts like ideas for possible games or products.
So I am not advocating no organization or tagging at all.
But I think a lot of people are unaware or at least haven’t really incorporated the distinction between information that is already present in the data, and information that must be manually added. So they spend a lot of time manually creating folder structures that encode information which could already be automatically derived from the data itself.
As for messages in closed apps, I just screenshot them. And I am relying on OCR technology to be or become good enough to refind those messages in the future. That way, if the platform itself is gone by then or the messages are not on the platform itself or hard to find on the platform itself for whatever reason.
So far I haven’t even needed to use OCR. Because if I look for a message I often have other memories of where I was, when it was or something else that happened around that time. So I just jump back in time in my photo stream and either find the screenshot right away or I find pictures near-by in time and spend a tiny amount of time looking forwards and/or backwards in time and I find the screenshot.
I do wish though, that iOS would automatically tag screenshots with the name of the app that the screenshot was taken in. And I think it would be cool if the screenshots were stored as SVG with pure text and vector shapes plus embedded bitmaps, so that the whole potentially needing robust OCR in the future thing could be side-stepped.
Your phone didn't recognize the object, you relied on a cloud service to do that for you.
When you use such services for free, you'll end up with all kinds of legal compromises that don't necessarily benefit you as an individual in the long run. Your personal convenience is subservient to other goals that don't necessarily align with public interests at large.
You could argue that the infrastructure will keep miniaturizing and one day you might not need those services. But that's not how things are currently evolving. Moreover, it will always take massive amounts of data to re-create the same models that are able to recognize patterns that are relevant to your specific context when you have a query.
At the end of the day, it's about what trade offs you are willing to accept. Cloud services based on machine learning do give you a good amount of convenience, but then you have to be willing to accept the hidden costs as well.
In iOS 13 (which I am running), it is indeed my phone that does this.
https://www.apple.com/ios/photos/pdf/Photos_Tech_Brief_Sept_... (page 4) says:
> Photos is enabled by powerful machine learning to deliver unique features like Memories, Search Suggestions, and For You. Photos analyzes every photo in a user’s photo library using on-device machine learning that delivers a personalized experience for each user. And this analysis is designed from the ground up with privacy in mind, with all of the processing done on device—and the results of this analysis are not shared with anyone, not even Apple.
> Photos uses on-device processing to analyze each photo and video in a number of ways, including:
> • Scene classification
> Identifies objects, like an airplane or a bike, and scenes, like a cityscape or a zoo, that visually appear in a photo, using a multilabel network with over a thousand classes.
> [...]
The training itself as you point out, happens not on the local device but on the servers that Apple own. So that part you are right about, but that is to be expected. Otherwise, manual tagging by each individual user (as well as significantly more processing power) would be required after all in order to train the models in the first place.
Naturally, "civilizational scale archival" is only feasible for a proper archival organization such as a museum, a library or an archive. As a person, you can't have this. You can use the archival-grade media like M-Disc, but don't expect to put something on it and recover it 50 years later easily. You have to design the process to validate and migrate the data every once in a while. Digital storage can't offer something comparable to a simple printed photo.
> when the vast majority of our knowledge ... would be unrecoverable just 3 decades after a global calamity.
The vast majority of our knowledge is encoded in the societal and economic context. There's simply no way to translate it to any media, and any disruption would be the end of it.
The challenge is keeping things accessible. Only copying does that for electronic media. That includes getting things like photos printed.
I think this is an excellent place for neural networks. They can preserve vast amounts of data compactly for many data types because they statistically compress high level abstract data which can then be used to fill in regions with high error rate, although if you did that at a large scale you'd probably end up with some constant error rate fluctuating around the average true value.
All indications point to the fact that we seem to be working against the unstoppable Force of entropy - indefinite error free preservation of data is ultimately impossible.
The idea is to act as a repeater, or like dynamically refreshed RAM, detecting and amplifying the signal at each copy before it degrades too far. You still have the error rate of reading and writing each time, but it's often cheaper to get the same error rate by refreshing cheap media than by using a more permanent medium.
Interleaved error correction (such as on a CD) kind of works as you describe, spreading out the error from a single physical point across the logical data so that it can be corrected by intact data elsewhere. This works because it's designed to recover from burst errors, corresponding to a localized scratch across the track.
https://en.wikipedia.org/wiki/Cross-interleaved_Reed%E2%80%9...
I recall working with a group of archivists where one camp was in favor of conversion to standard formats like PDF for certain documents, another wanted to preserve documents in the original format, and still another wanted to do both.
It’s a hard problem and purity of thought makes it worse.
I am not sure I understand. Isn’t the error rate zero? And wouldn’t you be using checksums to verify perfect copies?
As a side effect, this made me increase how much redundancy I used for data I really cared about.
I’d be more worried about disasters and budget cuts than cosmic radiation.
[1] https://recon.cx/2015/slides/recon2015-19-mike-ryan-john-mcm...
The second choice is using hard drives (easily available) and every so often power them up and copy data to new drives.
If you have a small quantity of data, then encode and laser print onto paper, with a font designed for optical scanning or QR code’s.
IIRC, the writers for M-DISCs are special, but the reader can be any DVD or Blu-ray drive.
Honestly, I think the biggest consideration for digital archival media isn't so much the longevity of the media, but the future availability of equipment to read it.
That's one of the biggest benefits of paper, IMHO. Besides being very well-understood material, nearly everyone is born with the necessary reading equipment and the decoding software is very common.
[Citation needed]
https://en.wikipedia.org/wiki/Literacy
There you go.
Because some people can’t read. Har har har :)
>The second choice is using hard drives (easily available) and every so often power them up and copy data to new drives.
I suppose the question is whether doing so could remain under the error correction threshold indefinitely, since there will be errors accumulating both during copying and over time in cold storage. If manufacture of new drives stops, it also isn't clear to me if only the data stored on them has a 30 year life or if the medium itself decays regardless of whether it is in use or not.
In theory I imaging keeping an unused NAND or even magnetic drive in cool dry storage should preserve it's physical integrity indefinitely...
TL;DR: The CD-R's didn't survive the process, but the M-discs did, with their data intact.
Are they still produced?
Thanks!
I'm sure I own 20 year old DVDs in same condition.
In addition, printed audio CDs are of a different build than CD-Rs which have been found to not be as resistant to moisture and light.
Mass produced CDs are produced using a different process than the one consumer CDR writers use, and they are thus a much more stable storage medium. The data-containing layer is literally formed out of metal using a kind of mold: https://en.wikipedia.org/wiki/Compact_Disc_manufacturing#Ele...
CDRs are written by altering a layer of dye with a laser, and that dye is very vulnerable to chemical breakdown: https://en.wikipedia.org/wiki/CD-R#Physical_characteristics
>A disc should always be handled by grasping its outer edges, center hole or center hub clamping area. Avoid flexing the disc, exposing it to direct sunlight, excessive heat and/or humidity, handle it only when being used and do not eat, drink and smoke near it. Discs should be stored in jewel cases rather than sleeves as cases do not contact the discs’ surfaces and generally provide better protection again scratches, dust, light and rapid humidity changes. Once placed in their cases discs can be further protected by keeping them in a closed box, drawer or cabinet. For long-term storage and archival situations it is advisable to follow manufacturer instructions. For further information consult the international standards for preserving optical media (ISO 18925:2002, Imaging materials — optical disc media — storage practices). [0]
---
Though, if you care about longevity, it might be better to use a technology like M-DISC. It uses a different recording technology to "[burn or etch] a permanent hole in the material, rather than changing the color of a dye."
It weakens. That's a shit explanation, I'd definitely look up "how does flash storage degrade".
https://archive.org/details/star-wars-ONTV-Early-80s/Star+Wa...
So this is more of a "oh yeah, that's what TV used to look like" moment than an actual "I want to watch the original Star Wars" again...
I have a friend at work who actually scored a pristine, never opened VHS copy and of course we ripped that sucker open and watched the day he got it. And we go back once in a while and watch it again, sometimes running it along side the new ones so we can spot and discuss the differences. Fun stuff on a Saturday night!
The true original is the one my parents taped off of TV when I was a kid, commercials and all.
What a brilliant idea! So obvious now, the teletext data must have been embedded into the tv picture thus captured by VHS recording.
This got me very excited, I'd love to rebuild the old Ceefax 101 pages.
EDIT: https://chrome.google.com/webstore/detail/archive-downloader... - It looks like that's a solid option to use a Chrome extension for it.
This is neither a rebuke nor pedantry - I genuinely hope this will be useful to you:
It's "en masse".
Also for some reason, Macaulay Culkin seems to be hanging out with them a lot. Maybe Milwaukie is just that much fun?
"Junka" -> https://youtu.be/9M39zY9OXFA
And oh yeah. Guilty myself.
There was a lot of interesting random stuff on them, mostly from the 80's. The bits of stuff he recorded is like a peek inside his brain :-)
Check out https://www.stitcher.com/podcast/behind-the-bastards/e/66732... for the details behind this truly unlikely person.
Kind of like Gutenberg.org texts, with their disclaimer that "we checked and couldn't find any renewal of copyright," or whatever it says nowadays...
However, it's such recordings that have saved many an old TV show as the studios reused tapes as well. Kinda how few Doctor Who episodes got saved.