Wikipedia and Internet Archive partner to fix 1M broken links on Wikipedia
blog.wikimedia.org
blog.wikimedia.org
Also, take a look at downloading a copy of Wikipedia! You can get a full download with images for around 100GB (last I checked, which was about two years ago). It's great for if you ever think you'll need some technical info while away from live internet. I keep a copy on a hard drive that boots into linux, just in case (maybe I'll be on site with a customer and need some engineering notes - BAM taken care of!)
[0] https://archive.org/donate/
EDIT: I used XOWA, and I do not keep the wiki up to date, really. Note that the entire wiki history is huge, but a reasonably current snapshot is manageable (~100GB or so).
I just set up a monthly donation amount, and I encourage you to do the same if you believe in the Internet Archive's mission!
If you think about it, that might be a game changer. Maybe it might make sense to ship e.g. smart phones with wikipedia, OSM and satellite photo (e.g. Sentinel-2 is free data) on disk.
Granted, 10m resolution is not the state-of-the-art in aerial imagery - which is around ~0.1m (this means 10^4 times more data) - but it is reasonable to detect buildings and combined with OSM's vector data you have practically Google Maps in your pocket.
I don't think this is a sensible use of local phone storage, but I do think you'll see a lot more P2P edge cache nodes if/when IPFS takes off (OSM tiles are already served on the IPFS network).
It would be a simple matter of picking a VPS provider or hardware colo provider near expected heavy use, launching IPFS, and having it pin the relevant content locally.
Perhaps, long from now, they might quote internet comments from sites like HN like we quote the Greek philosophers. "cmdrfred the elder once said...". I should make a habit of ripping the front page or so and comments for the ark.
Or Talmudic claims of authority in Judaism:
https://duckduckgo.com/?q=%22taught+in+the+name+of+rabbi%22&...
https://duckduckgo.com/?q=%22said+in+the+name+of+rabbi%22&ia...
Muslim ahadith also have the phenomenon of the chain of narration where people declare the provenance of the teaching:
The obvious roadblocks are:
- Websites, even content-focused ones, can be so complicated in terms of JS that the Archive might not capture it accurately.
- The Archive's policy is to honor robots.txt and other no-archive directives.
That is an advantage. It reduces extraneous incentives to post links, which should be only about the information they provide and not about page views. At wikipedia's scale, and sensitivity to information purity, that's relevant.
See also the rel="nofollow" decision from a few years back.
Regardless, neither the Internet Archive nor Wikimedia use AWS or other cloud providers, as it would be prohibitively expensive. They both run their own infrastructure/ops.
1. Get Wikipedia to send lots of requests to Archive.
2. Archive blocks requests from Wikipedia.
3. Wikipedia citations are disabled.
In other words getting Wikipedia to DDoS Archive, so that Archive's defense hurts Wikipedia.
A very silly scenario of course, just coming up with one for why an attacker might want to indirectly DDoS Archive via Wikipedia.
Though it's certainly possible to generate enough junk edit traffic to cause disruption on Wikipedia, but that's nothing new. It's the nature of Wikipedia as the resource - it trusts the internet community to be good on average. So far it worked.
The Wikiarchivator?
What bothers me is that they do so retroactively based on the current robots.txt, not the one contemporary with the archived content. So if a domain parker takes over a domain, and their robots.txt excludes everyone from every page (or everyone but Google), then archive.org no longer provides its archive of the old content.
Still, if the domain owner changes, they should not be able to remove content from old archives. That's like being able to remove stuff from encyclopedias about a palace somewhere, just because you live in the place where the palace once stood.
They are The Internet Archive after all, it's logical to archive contents like the domain owner and robots.txt for a given point in time. A change of owner can be easily detected.
Are you going to fund the internet archive to handle that workload?
> if the domain owner changes, they should not be able to remove content from old archives.
Why not? If that work belongs to anyone, it is the current domain owner. Why does the fact that the internet archive happened to crawl it mean that suddenly they lose control of their information?
This is elevating the Internet Archive from 'hey it's cool someone made a copy of that while it was up and no one cared' 'because the Internet Archive crawled it the world has absolute rights to that information from now on, wishes of the owner be damned.'
Not at all. A domain parker taking over a domain does not imply they have any rights over all the content that the previous owner of the domain posted.
I figured someone would ask that. I don't have an immediate answer but it's a good question (upvote for that). My hopes are just that removal requests are not too frequent. But without current numbers (of number of pages hidden after-the-fact and current removal requests) this is guesswork.
They could charge for it perhaps? A dollar per request. Doesn't seem too unreasonable for something you mistakenly made available to the planet. It doesn't have to be per page, so if you made a million documents available all under example.com/hidden/ then hiding that folder is a simple action and costs just one dollar. You're paying them for their time.
In the Netherlands, if you want your personal information (e.g. phone number; email address) removed from a company's systems, you can request that and they must grant it if they have no reason to keep the data any longer. And you can make requests to see your data, etc. But the law allows for companies to charge for this and I've seen example amounts (I think around 3 euros) somewhere. It's a somewhat similar situation.
So I don't have a single good answer, but I think by-case is a better way to go (and worth thinking about, at least) than just using the current approach.
If someone requests content be taken off, they instruct them to update their robots.txt. The content is not removed but will not be shown through archive.org as long as the robots.txt exclusion is in place.
There was a court case where the plaintiff wanted to subpoena the Internet Archive for evidence (since the defendant had since blocked the content with robots.txt). They sent an expert to testify that complying with that kind of thing would be too much of a burden for them, and suggested that the court force the defendant to change their robots.txt. The court agreed.
It would be perfectly possible to have the Wayback Machine respect the robots.txt and not have it crawl or archive any new pages, whilst making pages that have already been archived accessible unless a specific user agent has been denied.
Sadly, it requires a bit of human intervention there.
To take into account Javascript etc, one could also capture a png snapshot of the page.
One of the more curious cases was CSIRO (Australia's national science and research organisation) which seems to have not only deliberately purged a fair amount of data (Graham Turner's work specifically), but has a robots.txt in place which blocks archival by TIA. That strikes me as ... downright curious.
I'm thinking that could be something to have an Opposition Senator bring up in Senate Estimates.
404: http://www.csiro.au/en/Portals/Multimedia/CSIROpod/Growth-Li...
Now available: http://web.archive.org/web/20120508210658/http://www.csiro.a...
That's among the specific links which wasn't being served by TIA earlier.
dredmorbius@gmail.com if you happen to track that down.
I worked at The Archive for a little while and one of the projects I worked on was to unpack about 300 TB of crawl data from the defunct search company cuil.com. It was mostly decent quality data in a standard format and after some grinding, the whole thing was converted into warc files which the wayback machine could use to show the URLs. The end result was that about 60 billion URLs came "back onto the web".
During the work, I was looking at the stones rather than cathedral but after I left and thought about it in detail, it was very satisfying. I was reading the book "A Canticle For Leibowitz" at the time and the general theme of cycles of history was in my head. That dovetailed very well with the work I had done.
If you're interested, you can download the dumps over here https://archive.org/details/cuilcrawl
As for the book, I think if The Archive had novel for a totem, it'd be "A Canticle For Leibowitz". Very much affected my world view when I read it and this thread has just kindled my interest again. :)
There was an article by Brewster about preserving the Internet in the scientific American which put the average lifespan of a URL at around 40 days. The said article now 404s and the only way to get it is through the wayback machine. :)
http://web.archive.org/web/19970504212157/http://www.sciam.c...
Enclyclopedia Dramatica is generally not a reputable source of truth, being the site that it is, but while looking for some more information on archive.is mirroring of links from Wikipedia articles, I found an article on ED that I found interesting. It is heavily advocating one side of the story but at least it backs it up with some links, which is rather seldom on ED (most links on ED usually go to other pages on ED in my experience).
[1] https://en.wikipedia.org/wiki/List_of_Wikipedia_controversie...
For archive.org it is a known, established and trusted organization. It's actually has an office within walking distance from Wikimedia offices, AFAIK :) - not that it is very important, just an interesting fact. The point is there's no reputation problem. But for site that is less known, there is.
I understand the frustration of people about not being trusted, but that's how it works - trust needs to be earned. I don't see any way to it but for whoever runs the bot to talk to Wikipedia community and earn their trust. Name-calling won't exactly be helpful here like some do here in comments. Shady practices used by whoever wrote the bot like using tons of IPs and not identifying the bot properly also doesn't help. You can't be sneaky and complain there's no trust at the same time.
Websites going offline is a huge problem. For example, the now-famous thread from which sleepsort originated (on 4chan's /prog/ textboard) isn't archived anywhere: textboard threads are immortal, so nobody thought to archive any threads until dis.4chan.org went down for good.
Thankfully, some bright spark managed to save the sqlite databases for most of the boards on dis to the Internet Archive, so I was able to track down the thread eventually.
OT: I considered applying to the Internet Archive last time I was looking for work, but their office is too hard to commute to coming from the East Bay. :(
http://activehistory.ca/2013/06/myspace-is-cool-again-too-ba...
Unfortunately, the Internet Archive was only able to get the non-logged-in version of the site. All those loud, obnoxious profile pages users spent endless hours working on? We only have oral histories now to remember them.
[+] http://www.archiveteam.org/index.php?title=INTERNETARCHIVE.B...
I might be way off, but doesn't 1M seems like a low number for wikipedia size? What is that in percentage of total number of links? Does anyone know?
I wonder if they will publish a list of replaced links after the fact?
It would be far more reliable than depending on Internet Archive when it may not have the page archived and more likely the time of the archive would differ from the time it was referenced.
It would cost some more disk space and bandwidth, which of course is already pressuring them but in turn would greatly improve usability and reliability.
(I suppose Archive.org would be asked to take the content down)
There are fewer more noble pursuits than archiving the sum of human knowledge.