Discussing about the 4-color theorem recently proved, latest version of C compiler, difference between a vacuum tube amplifier and a solid-state amplifier, where GNU Project and Linux kernel was launched, and early online culture and tons of colorful, hilarious, but forgotten and buried memes, and weird phenomena emerged from the collective (un)consensus... Sci-Fi fandom being an integral part of online and hacker culture, millions of lewd story written in alt.sex, "Immediate Death of Usenet Predicted!", "There is no cabal", alt.french.captain.borg.borg.borg, Coffee and Cat warning, The church of Kibology, Anti-spam Movement, creationism vs evolutionism debates at talk.origin, Meow Wars - the first meme war online, all the personal attacks, trolls, flame wars, and "cyber-stalking", etc.
Then centralized WWW replaced distributed Usenet, crappy HTML replaced perfect machine-readable data format. Would we have a similar archive for Reddit or Hacker News? Possibly not. So Hacker News, just come and create one! You can make it! Anther unique challenge created by WWW is the inaccessibility of server-side software - exporting and preserving the data is NOT enough, unlike Usenet which you can just load any data. The user-interface and functionality of one website itself is also the collective memory that needs to be preserved - we need replicated software of a website, which has identical user-interface, which has all the functions from the original website: users to click an username and see the posts, karma of this user, etc. I don't think anyone even noticed the existence of this problem. Luckily, major websites online such as Reddit or 4chan, all use FLOSS software which would make the work easier, but still a huge challenge due to the inaccessibility of raw database. Also, to make some contents meaningful in the future, external resources such as hyperlinks to other websites and images should also be preserved, considering this, the chance of creating an authentic and complete archive is even lower.
---
But even if we're still using a distributed network where data preservation is still technically possible, and there is no walled garden, it may still be difficult to implement. In the era of Usenet, you often attach your name, address and phone number - there was virtually no threats except for a few trolls - this is why archiving Usenet was possible in the first place. But the Internet is not the Net anymore, now not only humans - almost every piece of equipment involved on the route may be your enemy.
The ongoing security and privacy movement is a huge threat of historical records. From my observation, at least of infosec hackers community - After Snowden's revelation, public and open discussion is slowing being transformed into private, closed, encrypted and temporary activity, plus self-hosted platforms like ActivityPub, GNU/Social, Mastodon. This is indeed good from a security and privacy perspective and it is exactly what we need now.
But we are also creating a huge gap of knowledge, information and history on the Internet. After my death, none of my self-hosted code, or my blog, or my GNU/Social posts will survive. In conclusion, "collect 'em all" is both an malicious NSA dragnet surveillance, and a glorified act of history preservation. This is where the contradiction lies.
I don't know what to do. For WWW, archive.org is a workaround and I think it needs more donation. But for all the other self-hosted things like Mastodon and git server, there is no solution at all.