Wayback Machine: Now with 240,000,000,000 URLs
blog.archive.org
blog.archive.org
The fact that they were able to preserve a masterpiece like this means a lot to me: http://web.archive.org/web/20010124071800/http://expage.com/...
[1] http://web.archive.org/web/19970414022225/http://www.scorche...
that's great ;-)
We were up to 20k daily uniques when I quit (not bad for 1997). I wrote the forum & related software myself in perl, which was an amazing learning experience.
That's exactly the kind of website they should be preserving! The future people will be glad that we kept some ephemera around.
And I love that website - just so blue; and comic sans not used ironically; and animated gifs; and the guestbook.
i realize this is only tangentially related, but your page compelled me to share. :)
robots.txt should have a limit; it shouldn't be applied retroactively so aggressively.
Reliable ownership change detection can be tricky though, but it's doable IMHO.
Hi, please I am using wayback-1.6 on my tomcat-5.28 (java-1.7 , ubuntu-11.04) to display all my arc.gz files but I have got this error, however this folder contains all my arc.gz files /tmp/wayback/files1/IA.arc.gz
Resource Not In Archive
The Resource you requested is not in this archive.
http://archive.org/details/10000000000000000BytesArchived?st...
Not even sure if archives can help - with some algorithmically created content it might be impossible to index it all.
Just one example - there are surely more. I used to think digital data would be easier to preserve for the future, but now I am not so sure anymore.
Not even mentioning Facebook, which presumably can not be archived because of the walled garden thing.
http://archive.org/web/researcher/cdx_file_format.php
Each CDX record maps a URL-timestamp pair to a byte offset into an ARC or WARC file. These are essentially just gzipped HTTP responses concatenated together:
http://archive.org/web/researcher/ArcFileFormat.php http://www.digitalpreservation.gov/formats/fdd/fdd000236.sht...
The document is retrieved, uncompressed, URLs are rewritten, the navigation banner javascript injected and the result is sent to the client.
The code is here: https://github.com/internetarchive/wayback