The Internet Archive and Jason Scott are saving our weird Internet history
blog.zamzar.com
blog.zamzar.com
I am a very pretty and flamboyant figure standing at the front of a literal army of real contributors to taking online history seriously and making it both available and accessible to future generations.
What do you think of the proposition that people should put important content on the archive on day one? ( Comment thread here: https://news.ycombinator.com/item?id=21843987 )
Also: Also thanks for the pennies a couple decades ago
Then there is the shared cultural knowledge at the time, such as Adam Curry (I think?) who on his own created MTV.com. Then later got sued and had to had the domain over when they realized there was value to it. On top of that, there was one of the first spoof pages I recall, "Madam Furry", which poked fun at Adam Curry's page. (Hopefully I'm remembering all these names correctly).
And lets not forget about Gopher, with Veronica and Archie, along with something else called WAIS (Wide Area Information Services), which always seem very slow and barely workable.
Oh, and how did we figure out the who/what/where? The Internet Yellow Pages, of course. Thick book that had everything categorized. I've still got mine around, brought it into work to put in the commons area book shelf.
WAIS has an interesting connection to the Internet Archive.
Long before the Internet Yellow Pages, we used Scott Yanoff's Internet services list. That's how I found out about the WWW. Also there was a big list of sites that allowed anonymous FTP.
https://gopher.floodgap.com/gopher/gw.lite?=gopher.floodgap....
Say Wikimedia would become dependent on the donations of a company, that company could technically force them to publish misleading information.
In regards to mozilla they are currently not financially independent.
Most of their money comes from google (in exchange google is the default search engine).
Of course, this is not to say that funding the Archive isn't also important.
Yes of course.
But both projects' biggest constraints are not (lack of) funding. It's bad mismanagement, or in the case of Mozilla, really bad mismanagement.
IE, this is as good a place as any to once again complain about ways that a significant amount of stuff enters a memory hole AFTER being put int to the internet archive (example ezboard.com but I think that's just an example[1]).
Basically, archive pulls content when a later robots.txt file says don't archive (the robots.txt file of the domain parker after the actual website closed, generally). Broadly, their approach seems like "anyone, anywhere who even implicitly claims copyright on X can knock it and anything related X off archive forever. [1]
And the last time I discussed this here, the policy itself had supposedly been updated but ezboard content in particular (the example in the link and something I'm interested in), still wasn't available.
[1] https://archive.org/post/389129/why-is-archived-content-purg...
Moreover, states and corporations are naturally going from incidentally targeting archived content to systematically targeting it - so that through a flood of incidental and purposeful copyright attacks, archived content may wind-up a very well "pruned" set indeed.
It seems like having a less "monolithic" approach to archiving may be necessary - I'm not sure what that would look like, networks of smaller archivers who try to cover everything? Tor? Freenet?
Edit: and yeah you can browse it in the Wayback Machine https://web.archive.org/web/20200113163734/https://encyclope... which has more frequent updates, but I don't think they're as thorough which means the inter-page links might go to versions from different dates. The dump files have a better chance of being good snapshots.
Do you enjoy this gruesomely inhumane shit?