Can the Internet Be Archived? (2015)
newyorker.com
newyorker.com
But more importantly if one thinks of Orwell's 1984, were all printed records were manipulated on a daily basis to reflect a different state of reality while the events were changing, in the days of internet that is not just a tedious sci-fi process but one that can be performed in a fairly efficient and methodical manner to an extent that it cannot be told what was real or not.
Seriously. It's a transient medium, somewhere between a beautifully illustrated book and a shopping list scrawled on the back of a napkin in crayon, recited down a megaphone by a drunk.
Let content live or die in its merit, rather than the desire to arbitrarily preserve it.
Not every book printed is worth reading, saving, and passing down. Not every thought is worth speaking and hearing. Not every meme, click bait article, or shitpost needs to be saved.
Honestly I think you're being extremely generous about how much signal there is in the noise. I think the decimal point needs to move at least two places to the right, but probably more like four.
Unless you want to unravel how Trump got elected?
The stuff people took the time to carve in stone survived. And not much else.
There are things that you can learn from discarded ephemera that won't get preserved in museums due to cultural values and taboos.
That's how you come away with the idea that ancient civilizations were perfect utopias with better, happier citizens.
I think the big problem here is that the durability of data is primarily based on a corporation surviving. Think of all that will be lost if Twitter goes bankrupt.
Plus the things that will be useful in the future might not seem useful now, so it may not seem worth preserving them.
This would be a great question to ask researchers in archaeology, anthropology, linguistics, history, etc.
My own impression is that the more data there is, the better for researchers. They can always weed out what they don't need later, but there's no way to get back information that has been destroyed or lost.
Also, in the past people have made some horrible choices regarding what's valuable and what's not. For all sorts of reasons, from political to psychological to social, they've done things like burn, discard, or destroy texts and artwork we now consider valuable but they did not. Often, what was considered valuable (like sacred texts) was not really as revealing about their daily lives as things they did not value and discarded.
So I'd really think twice before we discard any records. If there's some kind of serious pressing need to be selective (such as lack of space) that's one thing. But if keeping it more would not be a huge burden, I'm much more of an inclusionist than a deletionist.
Case in point:
We already know the answer to this one. Archaeologists are most excited when they find the village garbage dump. They learn far more about what people's lives were like by sifting through trash than by marveling at pyramids.
Archive the internet: crayons, megaphones, Twitter flamewars, and all. What's trivial to us may be critical to archaeologists of the future. They'll be far more interested in how we actually live than in how we want them to see us.
In an article about this preservation effort, there's an interesting couple of paragraphs:
We started dumping stuff that we thought was obviously of no future use, groups that specialized in a lot of talk and no substance, so to speak. For example, fairly early on there was a newsgroup about abortion which specialized in violent arguments.”
That’s why not only the very earliest Usenet posts, before Spencer started archiving in 1981 (Usenet began in 1979) but even some of the posts in the 1980s are still lost. It’s too bad; today, wouldn’t more of us rather see what was being said about abortion in 1984 than sift through the arcana of bug fixes in systems that have probably been long since retired? “It was perfectly reasonable from the viewpoint of stuff that we might want to use again, but a little sad from today’s viewpoint,” Spencer admits.
The downside of this, is you end up with buildings full of things no one looks at. You can't hold onto all the world's art, history, and scientific samples, without eventually throwing some stuff out.
It will probably never be cheap to store something like "The full set of all Facebook photos and videos", which means we'll probably end up keeping a sample of some of this stuff.
There are achaeological artifacts (and paleontological ones) which are old, but not particularly useful.
Price = Demand / Supply
The number of times per week I need to find a link that's gone dead—even occasionally for fairly recent articles—demonstrates the practical need for preserving web content, even just for pure contextual reasons. There's an enormous amount of useful content archived: information, original documentation, downloads, culture. I have a lot of respect for internet archivists.
A large portion of the Internet will be archived regardless. Not archiving all of it just allows a more distorted picture of what happened.
Any piece of information that's produced is not simply "content" having "merit" but evidence for the structure of this society as a whole. The most ridiculous content still sheds light on other content and gives an explanation for the origin of still other content and even events.
And, of course, even if information has a set merit, it won't necessarily survive or die on that merit but rather on a combination of merit, who it money for, purely random factors and so-forth.
Absolutely. Without a doubt. And not because the merit of a particular piece.
Those would be incredibly useful to have. (I even went so far as to find the owner of one of the defunct sites and tweeted him about the files. He said they were in a storage unit somewhere and he would get around to reuploading them "eventually.")
But content in every other medium has lived and died due to happenstance, religion or the aesthetic tastes of the wealthy. Why should the internet be an exception?
Or it doesn't even have to be profound... I was talking to someone yesterday who mentioned that their corporate site had actually been online since the early 90s, despite them being a non-tech company, and he wished he could show us the old designs. So we pulled up some old designs on archive.org. They certainly weren't archived for their historical value, but it was a nice experience.
I am working on PageDash (https://www.pagedash.com, just a Google form for now). The idea is to archive web pages for your private use via one-click via a browser (Chrome first) extension. So I think it jives with that people are saying here: archive what you think is important to you, leave the rest.
PageDash takes a different approach in that it doesn't use a browser in the cloud (usually PhantomJS) to archive the pages. This means that PageDash can archive when you are logged in, or when you are on a private network, for example. PageDash tries its best to preserve the page exactly as you saw it.
Sign up to be notified when it launches!
More importantly PageDash is web hosted and will allow users to share their archived links (private by default). Eventually, full-text search and shared folders can enable collaboration. Basically stuff that an offline-only solution can't do.
Ideally, PageDash turns itself over to be a non-profit sort of thing. Else, a corporation takes over. If really shutting down, there will be an export option for users for a graceful shutdown.
The value depends on the people who are trying to dig up information about their ancestors to find out how they managed to live in those primitive times.
Historians could always use more information from multiple sources to piece together accurate pictures of times past and the reasons why people did what they did. We know that since we have dedicated people trying to reconstruct our pasts. We'll have to live with what survived from what our ancestors left us but there's no reason to limit generations of future scientists to such scant and badly pieced together material.
But it would be much more fun to transcribe it.
You know, the way they used to copy books back in the days before the printing press.
Lock yourself in a cave with some (hemp) oil candles, paper and an iPad and ... you'd be doing humanity a huge favor.
When, in your monumental quest, you come across this message, know that somebody thought of this moment a long long time ago and hereby officially thanks you for your effort.
I needed to reference that list and spent hours scouring the internet for individual mentions of awards and placements to piece together a partial view of the results for one state for some of the years. It was horrible, but the worst part of it was the realization that this is but a single example, that the impermanence of the internet it's going to lead to a very sad loss of some very important data that we will dearly regret in the years to come.
It's also no longer sufficient to cache text and HTML; sites like NYT and WaPo have put massive work and countless man-hours into web apps that contain valuable data that relies on the presence of a back end server to populate the front-end, and rich JS apps to portray that data. It's going to be a challenge.
To the sibling comment -- as far as I know, no one's ever seen those tweets. You can get content back out of the Archive, at least.
It's common on HN for older stories to get posted and get upvoted to the front page, and commenters will make a note in the comments that it is an older story and should be marked as such in the hopes that the mods will notice and add e.g. (2015) to the title.
The tersest way to do this is to simply post a top level comment with the year the article was written, surrounded by parentheses.