The People Behind the Wayback Machine
motherjones.com
motherjones.com
I feel humanity will look back on this period -- lost diskettes, CDs, DVDs, game consoles, Betamax, VHS -- and lament the comparative black hole of information. Myspace history deleted with no warning, Justin.tv deleting history with only 8 days warning[2], ... and a million more examples. The Internet Archive are fighting that one bit at a time.
[1]: https://news.ycombinator.com/item?id=7826313
[2]: https://twitter.com/internetarchive/status/31565557291982848...
Archiving books, scientific journals and the likes would seem much more useful, but obviously you'd run into copyright issues.
Seems like a hard problem to solve. The low-hanging fruit would probably be detecting duplicates and combining them, which loses redundancy but handles all of those identical landing pages.
He said: "I guarantee that in the future researchers will curse us for having missed something absolutely critical. But only people using the archive can tell us about mistakes in what we collect. There is a cheaper alternative concept, called 'dark archiving', which means that we should not give people access to them. But preservation without access is dangerous - there's no way of reviewing what's in there."
But later on, he mentioned that: "AltaVista was the first Internet search engine that tried to be a complete index of all the pages. But what really got me was that they threw away the original pages. That grated, no end."
Aside: Kahle was one of the founders, with Danny Hillis, of Thinking Machines - the company that created the fabulous 'Connection Machine'.
But I know that the owner of days posterous page had no intent on keeping the page a going concern, and was happy to see it disposed of. In light of the recent Google "right to be forgotten" ruling, will there come a day when the right to be forgotten will extend to archive.org?
Sites can at any time opt out of being archived via a robots.txt exclusion (IA still keep their previous archives privately). However for public blogging sites operated by a third-party that's another matter.
Contacted the Archive Team about a few sites that would be otherwise lost and they archived them, and will eventually be progressively be uploaded to IA. Great folk there as well.
For instance, this broken piece of shit was my attempt to do something about the LMA interface -- http://www.archive-ui.org/#/. (select a show, reload the page, then it'll work. got busy and lost interest).
The interface is as bad as you say, but at least they give us the ability to do something about it ourselves.
It's something like the difference between your local used book store, and the Library of Congress. Or maybe the difference between the display cases of your local natural history museum, and the basement of the Smithsonian.
They web material was distributed for free in the first place. They're redistributing ad-ware, not stuff behind a paywall. (The same can be said of some TV shows and indeed I think TV show piracy if often met with a comparatively cavalier attitude.)
It's used as a measure of last resort. If I want to read an article from Wired, I'm going to try to find it on Wired -- or more likely, I'm going to Google it and get a link to Wired, and not the archive. It's only when it's unavailable from the original publisher or when I have specific historic interest that I end up using the web archive. The result is that publishers aren't denied their ad revenues as long as they host their material. Your abandonware argument translates neatly to the web archiving efforts.
They're archiving. This gives them a touch of academia and altruism that's casts them in a totally different light.
None of these things in isolation necessarily makes what the IA does entirely legit under current copyright law; they effectively operate in something of a legal grey area. But add it all together and not many people are going to get upset--especially given that they'll remove material if asked to do so.
There have been a few legal cases http://en.wikipedia.org/wiki/Wayback_Machine but not many considering the scope of what they archive.