A note to young folks: download the things you love
birchtree.me
birchtree.me
I recently built a new PC and, while I didn't migrate everything I saved, I made sure to keep my old system around for a few weeks just in case I needed anything.
Someone in a community I was in asked if someone had a link to some old Call of Duty shitpost from a decade ago, and I immediately recognised what they were on about. But a good hour spent scouring my early YouTube likes revealed nothing, so I dragged the old rig back in, powered it on, lo and behold within a minute I'd found the exact video I must've saved eons back.
No idea who created the said video, but I'm sure that if I noted a name or channel URL, there'd be a lot more missing media out there which others would be looking for.
--
On one hand, I respect the fact that YouTube facilitates a level of privacy when you delete or hide things, obfuscating any trace of the video from playlists, etc. But on the other hand, the amount of times I've known I've added a video to a playlist only for it to be deleted with no trace... Yeah, it can suck. Recently turning a lot of unlisted videos into private ones hasn't helped in that respect either.
A more robust method is also saving the full metadata to a separate JSON (though if one likes to rename files then it has to be remembered to rename both the files).
Given the rest of your comment, that seems like an interesting decision. Why not migrate everything?
I moved as much as I could onto some spare USB sticks and used them to pass the files along to the new build. I prioritised my photo archive (backed up online too) and project files first, then music, and then assorted videos and images. I prioritised the ones which spoke to me the most, which I came back to often, those which inspired nostalgia, and those which I felt would probably be gone in the next few years. It's a little short sighted, but I was strapped for storage.
The migration was only this month, and while I wiped my old SSD to use as an extra drive in this build, the harddrive which contained my old downloads folder is still in my possession... And only now is it coming to me that I can just plug it in to my current build and transfer the files
--
In terms of proper backups, I've yet to get something set up. My photo archives are backed up online and on multiple USB drives, and everything I migrated is on my new build and also on a series of USB drives... Again, I'll probably prioritise personal files moreso than random videos I downloaded, but it's given me food for thought.
Then streaming arrived. I stopped pirating as finally I could pay for legal access to a couple (prime, netflix) of decent media libraries.
Over the past few years, as the market has fragmented into a dozen £30 a month services, I started pirating again. I’m not spending £300 a month on streaming services, nor should anybody.
Then, they started pulling content from their libraries, or editing it to suit today’s sensitive sensibilities - and this cemented my conviction that what I was doing was justified.
Now that I have a vault stuffed full of UHD rips of a century of cinema, the other argument which has become self-apparent is quality. The UHD content streaming services serve looks like vegetable soup in comparison to a 70gb+ encode of a film, and is worsening as streaming services seek to be profitable.
I see another cycle of the same type of anti-MPAA/RIAA movement brewing as was seen around deCSS etc. last time - as they are once again going down the same route of making the only offering a crappy and overpriced offering, and the alternative is easier than it ever was before.
Now that Netflix is increasing its prices and kicking off account sharers while losing content to Paramount/Disney/Apple/whatever I've cancelled all subscriptions and I'm back to piracy/torrents.
I'm only paying for Crunchyroll now because hunting down Anime in good quality with working subtitles is too cumbersome. Unfortunately the app is terrible to use and unreliable as well but still slightly less so than the piracy sites.
The issue is not space. The issue is the hoarding itself, and the messiness that is inherent in that.
Move it all over
Never think about it again
what good is a pile of mediocrity where you can't find the gems inside?
What's needed is a wayback-machine-esque archiving method, where you can just find old stuff, but locally and self-hosted! Unfortunately, it does take a bit of effort to setup and not at all easy/painfree to keep running.
So i guess my box of burnable dvds will forever be relegated to be just a display item.
I'm saving up for a new tape drive (will probably be LTO 8) ... I have some fairly serious data hoarding issues.
Though I've not used recordable CDs/DVDs for much for years, just occasionally when posting digital photos/videos up to a low-tech type who has limited slow Internet. These days online drives are large enough that I can keep everything I care to have to hand on a local RAID array for easy access, backed up remotely of course, and Internet access is quick and reliable so things that are less important or rare can be streamed or otherwise re-obtained easily. Maybe quality has improved over time?
A few files seemed to read without error reported by the OS (though slowly and you could hear the drives retrying the read a few times) on more than one drive but had corruption from one of them, we manually compared different files of the same name to pick the best. Considering this, I'm surprised how much we did manage to read successfully.
There were other things on the disk, MP3s and other such, though we didn't bother trying to rescue these (only the photos were of any significance) so I'm not sure if they were in any better/worse state.
Mildly interesting observation: he no longer had a CD/DVD/other reader in a computer at all. No desktop PC and his laptop has no optical drive. The only optical drive he has is a DVD player hooked up to his TV and that is hardly ever used, as most of what it would be used for he has access to via a streaming service or two. I suspect this is not uncommon.
Unfortunately they need to be considered before you need to recover the data, which my friend did not.
Also/instead having multiple copies would have been useful as one may have survived better (stored in different locations in case the environment in one makes the disk degrade faster). Though neither method removes the need to occasionally test the media to make sure you have enough workable bits to recover the data.
For anyone who doesn’t know already: they got popular (and probably created) for binary files posted to usenet. Usenet was a high-speed and low-completion platform where files got corrupted all the time. PARity files took the algorithm for RAID and created a new file that could live “next to” the binary files. If you had a 1GB file and 300MB of PAR/PAR2 files could could recover the complete original file even if 30% was missing or corrupted. It still works today on Windows/macOS/Linux.
As a not to others: if following this method for archive purposes, be careful how you are distributing these parity files and those that they relate to. If all your files are on the same physical disk/tape the one physical disk/tape being unreadable kills the whole lot so you are only protecting yourself against errors in individual blocks. To protect against more complete media failures (i.e. a CD/DVD being unreadable) you need to split the data over several disks/tapes. This matches how the system was used on usenet, where each post (of potentially many tens or hundreds making up a large binary post) was essentially a separate bit of storage media. At this point it may just be easier to have multiple copies, though note that even with multiple copies the parity information can be useful as you can detect bit and block level corruption and in many cases correct it (whereas with two copies you can detect corruption, as the copies no longer match, but have no reliable way of knowing which copy, if any, is correct).
It's like physical hoarding: If everything's organised, it's a delight in years to come, but if it's in a big box-o-crap, then it's the same as not having it at all (worse, I guess, since it takes up space).
Cost of digital storage has just completely plummeted, and keeping stuff organized, even remotely, really makes for enjoyable rediscovery.
Even the mediocre stuff is nice to go through if it's sorted.
Old photos, videos and other original content is the best stuff to save for when you're older.
What helps for sure is like you mentioned in taking things at a more curated pace, where one can organize individual pieces of content at the time it's encountered (eg: saving a web page locally you want to refer/link back to in the future, or buying some music when one likes it and properly/consistently tagging it).
I have tens of thousands of web pages saved for example and it's a wonderful resource as search engines have become poorer for many results and sites/pages disappear. Due to the habits formed over time with filenaming/organization it's also super quick to find resources locally.
Modern day hoarders. If it wasn't for the Internet and hard drives, they'd fill their own space with stuff. :)
I got through everything, deleted about a quarter, and even took about 10% of my CDs to the recycling/exchange place.
This would be much less practical with video.
As someone with almost no working memory and fairly bad long-term memory, I'm so happy to have this sea of digital (and analog) detritus to sift through, to help me remember who I was and what I saw.
Everything I ever downloaded, inside "Desktop/Desktop/desktop/backups/oldpc/desktop/downloads/desktop/" awaits layers of my life, ready to be explored anew. Everything from the silly gifs I got on MSN messenger, to .pif files pointin into the void to remind me of that game I once played so often it earned a spot on my desktop.
Terabyte after terabyte of who I was and what I did, it is my most important treasure, and I enjoy strolling through it in it's pure, raw form.
I have recently realized there may be a couple legacy drives that might not have been wiped that might have it. I don't have much hope overall but I've been meaning to dig them out and check.
this type of thing is one of the undersung benefits vs a sea of ad-hoc hard drives. yes, a central server/NAS (even just 2 disks) is expensive but it's also a central point you can rely on and centralize your backups.
you can't let it become a pile either way, of course. but it's much much easier to work on manual dedup with everything on a single big zfs volume so you can compare directories and try to fold everything in sensibly. and snapshots let you roll this back if desired, or clone and copy files back out of history.
I seem to recall there was an algorithm called "downward flooding" or similar that was designed to help identify this directory-tree similarity problem. But I can't quite locate it.
It felt like it would last forever...
I think the issue is space though, otherwise we would simply archive everything and deal with the problem of search. For example, if you had every podcast audio file, you could search them using AI with something like "male and female talk about hacking floppy disks" and leave that to compute a while.
Then that fact sat in the back of my head for a bit and I wrote it off since “yes you could record it all, but then what? It takes so much time to filter and edit that you’d have to make the choice between filtering and actually living”.
But now, we’ve reached the edge of a different threshold: AI will be able to do the filtering and sorting.
I suspect that this means those who stored everything they were even remotely interested in will have an advantage. They will be able to have AI reveal the gems that others will have lost completely.
I hoard old scifi films and series (along with anything that I find of interest) and sometimes, the only copy that I can find has been posted to YouTube by some kind sole (as opposed to the copyright holder).
One particular service that really bugs me is the BBC iPlayer. The search is atrocious, so it's one of those sites that it's easier to do a search for "iplayer ..." in google etc. rather than using the BBC's own search. However, what really bugs me are the time restrictions and removal of programmes when as a TV license payer (which is how the BBC is funded), I've already paid towards the production of those programmes.
e.g. The excellent Oppenheimer series from the 1980s wasn't previously available on iPlayer despite being a BBC production and after the new film was released, the series was entirely available on YouTube, but not the BBC (they've since made it available).
Are you saying that people are using their feet to post videos to YouTube, or that there are actually fish that use computers now? If fish have become this intelligent, we're in trouble.
I always thought this was daft, if I had to choose between the BBC and one of Rupert Murdoch's offerings I wouldn't even have to think.
Couldn't agree more and would represent a real monetisation opportunity for the BBC even within the UK, if they are even legally allowed to do so.
How much would you pay per month for access to the BBC archives? I would pay quite a bit.
It doesn't even have to be streaming. Let's say you put in a ticket one day requesting that Oppenheimer show and a few days or weeks later, you get a temporary download link sent to you by email.
I started focusing my library on obscure titles and I make sure to keep my client going at all times. I've noticed my ratio jumping for the weird stuff lately. And there are also titles that I simply can't track down anymore because their seeders have all died.
It’s a shining example of public media done right and I wish more countries had film boards like it.
Build your digital life with this in mind, everything that ever matter to you in the slightest, must be stored on devices that are under your control.
I'm afraid of the future, I'm afraid of the day where I can no longer install Linux, and then, of the day where I can no longer build my own PC from parts, and the day when I can no longer buy a general purpose computer at all. When, not if, but when that day comes, all I have left are whatever cutting edge devices of the last generations I have in inventory, and I will be stuck in time, and I will be content with that, until my hardware supply dried out and I'm back into the world.. Hopefully I die first.
But this is expensive and cumbersome and I am aware that I'm a liability, if I die, my family have limited access to my data. I'm preparing instructions so that they can extract it, I'm keeping the family-relevant parts of the data on a read-only share that multiple family members can access.
Of course, if the server goes down while I die, little hope is left.
Printing out photos is definitely a good start.
And the cost is decent at $5/TB/mo (it has gone down over time).
The server can easily sync changes on a schedule and if there are concerns about privacy you can also encrypt before sending (add your decryption key to the instructions).
Of course, you should probably write up some instructions so they don't need to hire an expert, but in the worst case where you get hit by a bus before you write them, it should be possible to get to that data. The real problem is if you're using encrypted filesystems, and you didn't bother to write the root password anywhere.
However, if you need to find it again a second or third time, then put some effort into putting those things in a permanent and easy to access storage, maybe even physical.
This way you keep only things you really love, rather than a gigantic horde of mediocre things you thought was cool at the time but is not useful.
Nope, I just download everything that I really like. Storage is cheap.
But when choosing what to save or not, I do apply some criteria:
* Popular / obscure? Obscure stuff may disappear easily. Popular stuff will be online somewhere.
* Freely available, or paywalled content? (how easy to obtain)
* Recently released, or 'oldie'? See above.
* Storage space (+ download time!). I'll be much more critical about saving a multi-GB BluRay rip than some funny pics, a few-hundred KB pdf or game dump for 30y old console.
* An educated guess about whether I'll ever want to use it again.
So what I end up saving, tends to be on the smaller side, sometimes difficult-to-find-online, and with 'replay value' (broadly defined).
Watch once & forget movie -> delete. I don't want to devote brain cells (now, or in future) on stuff that isn't worth revisiting.
I found that the wayback machine is really useful to get the title of the video, especially for songs.
On most projects and topics i work on i keep personal notes. Sometimes i copy entire pieces of information from websites into my notes so that i don't lose it when the website dies. Any interesting webpage i at least archive through both archive.today and archive.org. Eventually i might look into taking that locally as well.
People have actually figured out a fancy acronym for it too: https://indieweb.org/POSSE
> POSSE is an abbreviation for Publish (on your) Own Site, Syndicate Elsewhere, the practice of posting content on your own site first, then publishing copies or sharing links to third parties (like social media silos) with original post links to provide viewers a path to directly interacting with your content.
Way too many (younger?) people expect social media platforms to be forever and keep that as the only place their content exists to a point where it's a part of their identity. And when that goes away, they practically disappear too - people don't have any alternative ways of contacting them.
This reminds me of a coworker I wanted to keep in contact with, but who only used Discord.
It was a funky(!) version of the Wizardry game tune.
One day youtube purged many “copyrighted” videos and somehow this marvellous version was disappeared forever.
I had been putting it off downloading it because I figured it was too obscure.
Aside from personal examples there is also all of the "right to repair", DRM, type stuff that gets taken down from YouTube, GitHub. There are exotic examples like TornadoCash (for AML rather than the above categories), but there are far more mundane examples of content just disappearing as well.
I've started a personal setup of mirroring the top starred repos and downloading youtube channels for content that I consider important (there is, of course, too much for me to capture it all).
It fits into my thinking about the Long Now. Working in formats that can be opened in 100 or 1000 years. I don't have the hubris to think my own content is particularly remarkable and worthy of persisting for 1000 years, if I find other peoples content interesting though I can help persist that for the future.
For my personal knowledge base I use Obsidian and link everything (including downloaded YouTube videos). I also use the Firefox extension "SingleFile" which allows the easy capture of web pages that you're reading (and supports a direct upload to github, which then publishes to vercel).
It's far from perfect but it's better than some of my early knowledge base efforts where I've come back and clicked on links only to find that they've disappeared (and in some case not even on the Wayback Machine).
This is a real problem for old stuff from decades ago, but much less so now: as long as the data is in some format that uses open-source code to read (i.e., open-source A/V codecs), you just have to make sure the source code is available for it. A lot of data from ages ago is really hard to decipher because it was stored in proprietary formats, and the specifications and code for those are lost or difficult to find. Data stored as PDF or h.264, for instance, probably won't have this problem.
Much of it might not be super relevant, e.g. there where a silly podcast called "We hate tech" hosted by two super inappropriate systems administrators, but I'd love to relisten, but it's gone. Videos from conferences who's websites are no long around.
Sadly I was never much of an archivist, but every so often I stumple on an old harddisk or thumbdrive with a small time capsule of my life many years ago.
Couldn't they just sell their image store to imgUr or something? I wonder if it's gone forever.
Imgur is no different, they purged images who were posted “anonymously” and anything they consider NSFW. They are very much part of the problem.
The relentless march towards obsolescence is hard to stop.
Yes, even those weird goofy projects. It's fun to go back and read it all again. I still have code I've written in high school and university, including a game that a friend and I implemented almost entirely inside a single C++ function.
I also now regularly use old code as a kind of mental bookmark about things I've implemented previously. When I'm working on a solution to something I'll sometimes have deja-vu, and I'm glad I still have the old code.
Me after downloading all of my DRM free purchases. With plenty of lovely names some going back to good old 8.3 character filenames... So in first place remember what you have and then try to map them like puzzle...
When we die, most of what we have will undoubtedly be considered detritus, our “junk”. However, our reflections on that material may be much more valuable, elucidating how we thought and how we were able to make some difference in the world; I know I personally enjoy reading my late grandfather’s writing about books and articles he read.
I think Tumblr was probably the closest we came to having some kind of media journal, but scribbling marginalia on printed Medium posts is a close second (joking mostly).
Personally I would lean towards SSDs mainly because an HDD failure can mean a broken mechanism and a nightmare recovery process while an SSD typically fails by becoming read-only. (Good-brand) thumb drives are also a solid option especially if you make multiple drives containing the same data and store some off-site.
Been debating getting a few and burning some of my more important stuff on a few of these:
Westworld RIP