Wayback machine gets a facelift, new features
archive.org
archive.org
and so on for infinite loop!
Oops this looks like recursion. ;-)
Should historians not read private letters sent long ago? Should they swear to some oath and take a moral stand that such things shouldn't be examined?
If the answer is "No, they should read them.", then in that same way, then why, for historical record, should we observe robots.txt? Isn't it the same thing?
Technically speaking a robots.txt that says
User-agent: * Disallow: /
means that you should not crawl the site today. It should have no effect whatsoever on displaying pages that WERE crawled before the timestamp on the robots.txt file.
I don't think there is an inarguable answer to my rhetorical question. People's intents and wishes do matter.
But there also is an idea from antiquity about the public good and the commons. I guess at some point my personal wishes get trumped by this overarching principle.
The whole point of the question was that someone would say "You may not read my love letters" and then society said "Too bad, we're doing it anyway. And reprinting it in highschool text books."
Is that ok? I don't think there's a clear line and I do think there are probably moral boundaries.
I'm by no means Lawrence Lessig and this type of discourse I'm really not experienced at. I do think there are many important questions here that we may need to rethink our thoughts on.
It's not our cultural tradition that every written work (train schedules, greeting cards, friendly notes, lolcats,etc.) must be archived at the Library of Congress. I'm not sure that it'd be a good idea.
Archive.org is a good idea.
To be more specific--short of granting the Internet Archive some sort of special library exemption--what if I were to say, create a special archive of popular cartoon strips. What's the distinction?
[EDIT: The retroactive robots.txt situation seems less clear but, like orphan works, also depends on the scenarios you care to devise.]
Also, the Supreme Court will be happy: http://www.nytimes.com/2013/09/24/us/politics/in-supreme-cou...
Bookmarklet (thanks to sp332):
javascript:void(open('//web.archive.org/save/'+encodeURI(document.location)))
document.location.href='//web.archive.org/save/'+$('#web_save_url').val();Does this look right?
javascript:void(open('//web.archive.org/save/'+encodeURI(document.location)))
I hacked it together from what you posted and what my archive.is bookmarklet specifies: javascript:void(open('http://archive.is/?run=1&url='+encodeURIComponent(document.l...)
EDIT: I can confirm that the bookmarklet I provided above does work. sp332, thanks for your help.
http://enwp.org/WP:Archive.is_RFC
http://enwp.org/WP:Archive.is_RFC/Rotlink_email_attempt <-- conversation with archive.is operator or representative
http://enwp.org/WP:Using_Archive.is <-- "corrected"
http://enwp.org/Archive.is <-- deleted as "non notable"
https://www.mediawiki.org/wiki/User:Kevin_Brown/ArchiveLinks
For various reasons this didn't get completed or deployed. It's still a good idea though. IMO it should be rewritten, but it wouldn't be a lot of code. I'd love to help anyone interested.
(French Wikipedia already does this, by the way. Check out the article on France, for example - all the footnotes have a secondary link to WikiWix. https://fr.wikipedia.org/wiki/France)
Could you list why? It looks like a sorely needed feature!
I would assume it's mostly that. They seem very accepting and willing to do a lot of things.
That's why I'm a "donation subscriber". If you'd like to know more about it, please visit: http://archive.org/donate/ - a subscription helps extra much, because it's a constant flow of cash. But one-time donations are of course of help as well.
The GSoC student didn't follow up with the process of getting it adopted. I didn't either, which I regret. I left the WMF in early 2012 so I guess it was dropped on the floor for a while.
That said I have since found out that others have taken up the charge.
https://github.com/internetarchive/wayback/tree/master/wayba...
(The CDX API linked below is links to the actual warc/arc archive files, not the web-viewable versions.)
Here's some info on the Memento API that links to web-viewable versions of a given url: http://ws-dl.blogspot.com/2013/07/2013-07-15-wayback-machine...
http://web.archive.org/web/timemap/link/{URI} will return a text stream of urls and dates.
Edit: formatting
Disclaimer: I'm really not trying to over-market myself, but I figured readers of this thread might be interested in my project. Happy to take down this post if it's read as too spammy.
http://www.archiveteam.org/index.php?title=Wget_with_WARC_ou...
http://www.archiveteam.org/index.php?title=The_WARC_Ecosyste...
Once you have the WARCs you can upload them to Archive.org and they can be added to the wayback, or you can set up your own service for browsing them, built off something like warc-proxy https://github.com/alard/warc-proxy (Yeah, same name different purpose...)
There is also a MITM version of WARCProxy that will let you store HTTPS sites: https://github.com/odie5533/WarcMITMProxy
http://www.archiveteam.org/index.php?title=Wget_with_WARC_ou...
This makes creating a browse-able mirror of a site in warc format fairly straightforward, as wget will automatically make links relative, as well as fetch requisite files (css, js, images) for each page.
They really should look at the date on the robots.txt and only apply it to pages retrieved while it is in effect.
Show us the pages from before the robots.txt became so restrictive!
But it would take hours or days to add every article from every issue.
I love this service.
Be sure to send them some!