The Ruins of Dead Social Networks (2011)
theatlantic.com
theatlantic.com
* There is an enormous quantity of culture locked into Facebook posts
* Since Facebook's graph API have been shut down, there exists no API-based way to export _other people's data_. You can get whatever you personally have posted to FB, but none of what _other people_ have posted.
* Individual posts might be scraped, and present in archive.org ,but people's / pages' indices are not available for non-logged-in users
* Social networks have a typical lifecycle of 5-10 years; and their shutdown can happen in a very short timeframe, see eg. geocities ; but imagine that without any accessible index, making full archiving impossible
* This implies, that if/when FB goes down, so will a major slice of Internet Culture circa 2008-2016
* If you have any suggestions, or partial solutions to this, please kindly post it to http://softwarerecs.stackexchange.com/questions/37036/self-h... .
I'm hoping these sites have archives, and that those archives would be made publicly available in case of such a disaster. But maybe even then they'll be buried behind a shitty web interface such as Google's web interface to old Usenet archives. Or maybe they'll just disappear completely, which would be a huge waste.
I've seen lots of omissions in archive.org. Not that it's not great. I love it. But there's a lot more that needs to be archived than what's in there.
[1] http://www.archiveteam.org/index.php?title=Main_Page
[2] http://www.archiveteam.org/index.php?title=ArchiveTeam_Warri...
On the flip side, a lot of the "information" on Reddit & al is just junk, about same value as pure noise.
That's the kind of common sense and succinctness I like to see. It's Yahoo so mismanagement or disappearance is a high risk. It's been acquired. Assets often get gutted or shelved after that. Two statements in five words tell you everything you need to know about why Yahoo Answers might need a mirror. :)
Note: I'm ignoring the complexity that copyright law would introduce into this to just focus on keeping the key info.
The best thing really is to do a wholesale archive of the whole site, and keep doing so at regular intervals. Just like archive.org is doing (at least for the sites it archives).
I think the digital form of the artifacts left over from social networks and the vast amounts of data retained is something we've not seen before and don't know how to catalog just yet.
This new historian job that can handle large amounts of unstructured data sounds perfect for a machine.
An excellent talk on this subject from 2010.