Firefox’s latest Test Pilot: No More 404s
ghacks.net
ghacks.net
If all it does is offer a link to the Wayback Machine, then it sounds great. But I worry that folks might want it to do more, without realising what that might mean.
No More 404s Telemetry ping is only gathering data about how often the Add-on is fired and how often it is clicked. We (Mozilla) don't know anything about URLs. The Wayback Machine has it's own privacy protections in place that you can learn about here: https://blog.archive.org/2013/10/25/reader-privacy-at-the-in...
Also worth noting, every Test Pilot experiment comes with a brief explainer of all data collections. Here's the explainer for No More 404s:
In addition to the data collected by all Test Pilot experiments, here are the key things you should know about what is happening when you use No More 404s:
* We collect basic usage on how many times you encounter a Page Not Found error (code 404), how many times a cached version of that page exists from Archive.org, and how many times you choose to view the cached version.
* To provide cached versions of pages, we send 404 error page URLs to Archive.org. Archive.org discloses its privacy policy here (https://archive.org/about/terms.php).
* We do not collect URLs of the pages you request or the URLs we send to Archive.org.
* We may share survey results you submit to us and aggregated telemetry data related to this experiment with the Internet Archive.
Why not only do that on a user action, instead of on every single 404?
The Archive is a lot bigger than just than the Wayback Machine. The collections include books, audio, video, and software.
There is an unofficial project by volunteers to make a geographically-distributed backup of the Archive. http://www.archiveteam.org/index.php?title=INTERNETARCHIVE.B... The basic tech is functioning now, you can run it http://www.archiveteam.org/index.php?title=INTERNETARCHIVE.B... but we're looking for developers to make it more user-friendly. If you'd like to help out, stop by #internetarchive.bak on efnet. http://chat.efnet.org
Is there really an upload cap? How can I upload faster so it is more feasible?
Thank you
To offer the user clearly explained options after a clear 404 does sound like a good idea however, granted that the Wayback Machine can (and wants to) handle the additional load.
> The notification reads: "This page appears to be missing. View a saved version courtesy of the Wayback Machine". You may click on the link to open the Internet Archive website to read an archived snapshot of the page on the site.
I started a Firefox extension and got the basic functionality working before getting bored reimplenting something I had already done. But now I've switched to Firefox for Android I'm thinking of reviving it, so I can have my plugin on mobile too. It's been pretty useful. I use it a couple times a week usually with about a 75% success rate at getting a cached or Wayback Machine page of a page that is either down or no longer has the content that used to be there.
Edit: Also, it only turns on per-tab instead of browser-wide which I could see annoyed everybody that used similar plugins for Chrome that would turn the whole browser into "browse cache-mode" when activated.
I don't. Can you explain?
Over the time scales that archive.org holds on to data, domain ownership itself becomes part of the history. While permitting someone to hide a mistake for security reasons is reasonable, allowing erasure of past owners' history by the current owner is counter to their stated purpose.
Given the prevalence bogus WHOIS data, the inverse is also possible: if the 2nd owner uses the same registrar and "privacy-protection" feature as the original owner, the WHOIS data could appear to have not changed, except for the start date of the registration, which would look identical to a single owner who re-registered their domain after allowing it to lapse.
As to the amount of traffic it generates, that's one reason why we're doing this through Test Pilot -- we can dial that up or down.
I don't think this is a great idea, it would make me worried about the Wayback Machine.
No doubt; I'm not sure what you're describing here.
So, I User put my user/pass into a form field and hit submit. The browser makes a post to site.com/login with my information, but the server returns a 404. What happens in this case?
There's no magic here; it's just a link that produces a GET request when you click it. I feel like a lot of the concerns and objections raised in this comment thread originate in not having taken a minute to find out what this extension actually does.
You can find more upcoming projects and information on Test Pilot here: https://testpilot.firefox.com/
When I have a mistake, the only thing I want is my UNALTERED original input, and a chance to fix it.
The ire-inducing thing about most “helper” pages, including the ad-ridden search pages that ISPs like to return instead of 404, is that they complete rewrite my address bar. This means that instead of being able to fix a single character and try again, I most likely have to retype the whole damned thing.
Especially on mobile, there’s nothing quite like entering "simplething.com/foo", accidentally typing "simplethin.com/foo", and being redirected to "godawfulisp.com/applications/helper.jsp?unnecessarycrap=1&adtracker=2&garbage=xyz&useextraobnoxiousads=true". At that point, my options are “go back” (returning a blank page) or trying to “select” the URL field (which is horrible on the iPhone) and reentering everything. All because I saw a page I didn’t even want. This made me so mad that I configured special blocker patterns to ensure that ISP pages could not even be loaded.
In other words, please stop “helping” users if you don’t really understand what the problem is. Sometimes an error and the unaltered original input is exactly the right thing to show.
The problem on the WWW is that you don't always get a 404. Sometimes the original web site goes up for sale and you end up with something like this: http://www.strictlybowhunting.com/Anov01issue/crows.htm which is not a 404. It would have to recognize that as being equivalent to a 404 and serve up the archive.org page anyway.
Ideally, a URL should allow the option of including a date.
Nothing spells out "fuck you" quite as clearly as breaking thousands of hyperlinks overnight.
On the one hand, we wouldn’t expect our daily interactions to be recorded and stored for long. Most of them, anyway – and certainly so for our most personal ones. On the other hand, the web is a medium where things tend to last a while, if not for very long thanks to efforts such as the Internet Archive’s. The delicacy of where to rest the equilibrium, I think, stems from the particularity of the web as a medium itself: a web page is kind of like ye olde book, but it adds the possibility of easy changes and user interactivity. In other words, if so used, a web will sit between traditional written media and face-to-face (or technologically mediated, but unrecorded) interaction.
Perhaps by custom, perhaps by nature, we expect personal interaction to be ephemeral, and we have no problem with other media, such as books, newspapers, etc., to store their contents for long. It’s no harm to publish a book with our thoughts and have it found across libraries in the world (indeed, it is a very nice thing), but we would feel violated if our conversations at home where recorded without proper cause.
The web stays somewhere in between. It’s a place for words to lie and last, but also for personal interaction. Personally, I favor ephemerality on the Internet, but don’t see it as characteristic of it. It’s not a given, and we should find a place for it, but only as far as personal interactions go. I want to be able to remove (or, at least, anonymize) whatever I post online in the manner of a casual conversation. But I want other kinds of content, such as manuals, news, etc., to be preserved from ephemerality.
I’m not sure whether The Internet Archive stands within these bounds. I would like to see it more as an opt-in facility for website administrators, instead of requiring manual forbidding.
Are there any plans to integrate, or allow integration of other services? If you could search for a dead URL on Wayback, Archive.is, Google Cache, and some others, you'd have a lot of coverage, and in the case of the last two, be immunte to IA's (frankly boneheaded) robots.txt policy.
It's an opt-in experiment, and at the end it's possible to a) be rolled in as a feature, b) spun off as an optional add-on, c) scrapped entirely.
Oh yeah, or d) taken back to the drawing board & re-engineered for another go-round.
The idea that it's appropriate as a plugin is in anticipation that they'd decide to integrate it like hello and pocket.
It's a feeling, there are all sorts of different browser efforts, FF just feels like we're getting to a point where it is drifting closer to the other browsers and getting heavier. YMMV.