Please, My Digital Archive. It’s Very Sick
laphamsquarterly.org
laphamsquarterly.org
It's nice to dive in the past from time to time but I don't need a 1:1 archive on 1996 online culture. There is so much low quality content dumped every second you'd be hard pressed to archive everything, youtube alone is hundreds of hours of video uploaded every minute, 24/7/365 for years.
Weirdly enough the only pics left from my childhood aren't digital, they're from a 1980s film slr (that I still use to this day).
Digital media is way more vulnerable if you don't know how to take care of it (taking care of physical photos is pretty straightforward in comparison). There's going to be a lot of people ignorant of the fact that media like CDs or even hard drives degrade, that lose their phones and the pictures in it without a backup, etc. whose memories will be wiped away.
In my country for example, there used to be a fb clone that was all the rage back when I was a teenager (since fb wasn't translated to our local language yet back then). The company pivoted and its social media site went offline after a few years, and many many people lost memories that were exclusively kept there - I met the CEO in a conference a few months ago and he mentioned that he's off social media entirely because of the amount of people constantly contacting him to get their photos back.
It's aligned with the Internet Archive in many ways and there is cooperation. They are looking for volunteers!
Other than that, if you design/build websites you could probably make sure that they are accessible for the crawlers and that there is technological compatibility. E.g. I've seen Internet Archive have problems with SPAs.
If you don't want things to stick around don't put them online.
I can understand overriding a corporate decision, but the meta tag on Mastodon profiles is a user preference that is explicitly opt in.
There's a lot of harm that can be done with archived data and you idiots never see the issue because you're not targets and you assume anything anyone puts on the internet will be curl'd by someone. But that's because YOU'RE THE ONES DOING IT. STOP.
It never was a point. It was common sense since the old days to never expect privacy from what you publish on the internet. That's what publishing means.
> But that's because YOU'RE THE ONES DOING IT. STOP.
No it's because the people who are doing harm using that data are doing harm. And they're not listening.
Keep in mind that archiving Mastodon is pretty much inevitable: spin up an instance, connect widely, keep toots.
There've been some announced archive projects, one of which flared out loudly a few weeks ago when I asked what the content removal policy was (the operator threatened retaliatory actions just for my asking, I saved their replies ... via the Internet Archive).
IA basically have an "ask and we'll remove" policy for content removal. Given that Mastodon toots are parented (at least for source instances) under the user's profile URL, asking for a global removal is fairly straightforward. Email info@archive.org
The Internet Archive's Wayback Machine is also, at least for now, not very easily searchable. Having been involved in the project to archive Google+, it's 1) impressive at what they saved and 2) disappointing at what was missed, as well as 3) frustratingly difficult to track down specific items. If I've got a link, I can generally get there. If not ... it's much harder.
I do think we need to come to some level of agreement/understanding of what can be / should be archived, and what should not. That's not easy. I'm not an archive absolutist (save everything), but I understand the value in saving even the ephemeral and non-obviously significant. Significance often increases in time and context.
Archiveteam, which my comment references, has been doing it against user wishes.
https://www.archiveteam.org/index.php?title=Mastodon
https://www.archiveteam.org/index.php?title=robots.txt
They're not the same entity, although people often assume they are.
They've been sensitive to exclusion requests, and serve principally to move data to IA, where, again, removal requests are easy and recognised.
IA have generally the same policy as AT on robots.txt, for entirely justifiable reasons.
In the grand scheme of things, open, public, accessible archives >>> closed, private, nonaccessible (or accessible-only-for-a-(price|class) archives. Adding the capacity to denote specific exclusions assists further in this.
I've lost my digital collections a number of times, and I just start over again.
I also regularly delete my old content/comments wherever I can.
Personal opinion, obviously.
Have you looked in an old paper archive? Hundreds of personal letters from Isaac Newton or Benjamin Franklin. They talk about all kinds of private matters, yet are very important for historical research.
Would you be happy with a future digital archive to have all your private chat messages to other people?