Make your 404s into 302s
4042302.org
4042302.org
But man, you really have to explain how it works a bit better. At first I thought that we should redirect 404s to your website and I was: "??".
What I understood: Which each iteration of the website, you archive the old one on a specific subdomain. Then, you redirect all 404s of the new website to the old one. Like that: no link is broken.
4042302.org definitely needs better explanations.
https://chrome.google.com/webstore/detail/wayback-machine/fp...
I use it, it's really useful.
> Here’s the simplest solution I could come up with:
> 1. Serve the current site from a subdomain (e.g., 2017.ar.al)
> 2. Make my 404s into 302s that point to the previous version of the site.
> 3. If I change the site again in the future, rinse and repeat.
> I call the technique 404 to 302.
In a sense this information conveyed by the 404 page is now the immutable 'resource' that will stay permanently at that URL. Doing a redirect breaks this, it's lossy and usually a bad UX.
> www → 2017 → 1997 → Show 404 error
edit: maybe not, because the old site was written to assume it's running at the current site's subdomain? i guess it depends on much you've changed your URL structure since then. that thought makes me a little squeemish.
it seems like a nice approach would be to return the 404, and make your 404 page render a link that says "try an archived version?". you gotta let the user know that what they're about to see might be stale.
Also, not sure if anyone is thinking of this, but there are security concerns if you were to serve up old pages on a new domain. If that old page has a vulnerability, it now has access to data in the new domain. This is more for archiving, so the old pages wont get patched.
In my travels, most folks who actually care about this (SEO links) build an extensive alias/redirect map for the old/previous core URLs they want to remain functional as part of their rebrand, when the new site goes live the redirects are dropped in as well. This is especially true if they've ever published the URLs on physical media (mailed postcards, e.g.).
Maybe the monkey-do approach to web publishing that underlies those decisions is the problem.
People who care enough about preserving history and current links probably already do that. People who don't care aren't going to start now because of this page. Especially those who have dynamic content and probably don't want to keep running a million different versions of their backend forever.
If you like something on the web then make a copy.
As a webmaster, if it's at all possible to go static (whatever your flavor of that is), then do that. A static website is easy to host and keep forever, and it's usually easy for consumers to archive too.
Not to hate on PHP, but keeping older PHP sites around securely has become a major undertaking. You can't safely run a wordpress site that hasn't been updated in 5 years because your security vulnerabilities are exposed to the wide web. If your static site generator has security flaws.. well that doesn't affect your current build artifacts and you can still run the thing in secure ways.
It was written toward the end of the BOHF's reign, when a technical specialist of the web had quite a lot of sway, when their decisions about a site's information architecture and how it was run was, if not the law, at least a very heavy hand on the till.
Those days are long past. Now Ted in Marketing wants a URL and who are you not to give it to him? I remember the pain of creating vanity top-level URLs in SharePoint 2003 because some functionary wanted them, and then they would promptly forget what they demanded. Yes, I used to use 410 Gones where appropriate.
That sort of thing has not been in our hands for quite a white, even if it is probably the best thing to do. After all, has the product URL changed? Or will it be back? Or has it been discontinued? The correct HTTP response, properly and widely used, would be very helpful in moving so much of the web forward but that is not under our control. Hasn't been for a while.
First, it unduly burdens the server, in sending multiple redirects to cover the entire search space of possible versions of a URL - for a mature site, this could be a lot of redirects. It also unduly burdens the client, in following them, and the network between the two.
Second, 302 is the wrong type of redirect to use here, because it is temporary; a well-tempered user agent will treat it as such, necessitating the same cascade of redirects on followup visits. The right way to do this is with a 301, which has a semantic of permanency, and is treated as such by user agents. But it's still the wrong thing to do.
Maintaining access to older versions of websites is, again, an entirely desirable thing to do. But if you're going to do it in a way that requires work on the server (as this design also does), you're better off just having the server maintain version information and serve the latest available page at a given URL, in a 200 response, when the URL is accessed.
from https://4042302.org/how/ :
>We use a 302 and not a 301 (permanent redirect) because we want the latest site to have the chance to override the URL in the future.
By filling up your sites with nested 302s (following this to it's conclusion, in ~10 years time), is not only a management headache, but may fall foul of Google (I'm not sure nested 302s, will send positive signals to Google) and result in your whole site being de-indexed.
But for a lot of medium sized companies with dynamic websites, this isn't always practical. They may not have the know how to dump their 2000s drupal install to static HTML files, and don't have the IT staff to upgrade and secure it.
I think there's always been a lot to be said for good, static sites, and there still is.
Re: Google and PageRank, pretty certain they've addressed this and recognize 302 and 301s and treat them the same. Previously, this was an issue.
The actual solution is put in the work and redirect the old missing page to a relevant new one.
If I had a link from vogue the BBC etc back in 1996 point to a product page to want to redirect that now broken link with a 301.
1. Create a 404.php (or whatever your preferred back-end is)
2. mod_rewrite real 404s to serve that script.
3. In that script have a lookup table/db/file that lists all the redirects you need.
4. Extract the requested URL from the server variables.
5. Use the lookup table to find the correct URL, and issue a 302 for it to the users browser.
It's kinda seamless, and I've been doing it for years.> You know why the web is broken? Because nobody cares.
I agree with this.
404s aren't a technical problem, they are a maintenance problem. If there was time, budget or interest to fix it, the 404 wouldn't have existed in the first place.
But I don’t think fixing broken links is at the top of our priority. Many links need to be broken. And I’m glad my old stuff isn’t around. Hopefully much of my new stuff will go away too. I think this is just humanity’s process - to sift through it all and hang on to the pieces people want to save.
You can recover a lot of value from recovering old links that would take a lot of effort to replace.
Doesn't that assume that all things on the internet are permanent? Why should that be the case? If I have a page on my website and I decide to delete it then I should be able to do that. Having links that pointed to it return a 404 is correct. 404s are useful. They convey real information.
Sites that do things like 302 redirection to the home page when the link is apparently incorrect are annoying - you can never tell if the page is really gone or if the website has incorrectly bounced you to the home page.
> Sites that do things like 302 redirection to the home page when the link is apparently incorrect are annoying
You seem to be operating on the assumption that this thing is doing something that it doesn't do.
>Indicates that the resource requested is no longer available and will not be available again. This should be used when a resource has been intentionally removed and the resource should be purged.
Still no. That kind of thing calls for a 5xx error.
Instead of 404s, you redirect to a previous version of your webserver (that is still running), which then instead of 404s redirects to a previous version of your site (that is still running), and eventually it tries the wayback machine.
So never pull down a previous version of your website, and issue 302s to that?
The site is suggesting a best practice to 302 FOUND-redirect you from:
<version x>.site.com
to: <version x-1>.site.com
Until it goes beyond the oldest version in which case you end up with a 404.At first I thought that it was simply some advocacy for taking the time to route legacy content using 302, but it also seems to be some sub-domain trick with years…?
It's really confusing to me.
This is not realistic for large SaaS apps. I run a SaaS app that has been online for over a decade, with millions of public documents. While we do strive to minimize URLs changes, and we do have robust redirects in place for old URL formats... it is simply not feasible to keep old versions online.
A REST request on a IoT device can return 404 if the device is not available at the moment. Any redirect is meaningless and actually breaks the semantics of 404. My understanding is that 404 has been so connected with the file like persistent resources, that people forgot the elegance of the wording in the standard.
I can understand the popular usage change, but I think then the standard actually becomes incomplete because it only fits web pages or document type resources when an HTTP status code is about Hyper text Transfer Protocol.
Alternate view disagreeing with what postulated above is that, actually HTTP has been abused and is being used in scenarios that have nothing to do with Hyper Text, looking at you JSON REST :D
Ultimately, HTTP has a very weak set of status codes for application layer concerns. The vast majority of the ones that sound good mean something about the HTTP layer rather than the application layer -- for example, it makes sense to return "precondition failed" if you request deletion of a directory that isn't empty (the precondition for deleting directories is that they are empty), but "412 Precondition Failed" means that the client supplied an HTTP precondition like "If-Match: abc123" and it failed. Trap for the unwary.
For this reason, I think REST largely fails to provide what API authors and users desire. If your API server can successfully convey an error to the client, you might as well say '200 OK {"error": "The Raspberry Pi you were looking for exists, but it's turned off or someone used its Ethernet cable to test their scissors and the test went very well."}' Now the client actually knows what's going on.
A better solution would be an intermediary part that never changes — say, CloudFlare — that caches HTML pages forever, automatically adds a “Archived content” header to the page, and warns the author so that they can either allow the archived version or make it a 404/410 instead.
Nobody wants to maintain servers forever, but serving static/frozen pages is much easier and cheaper.
Due to laziness and inertia I've kinda just left it there as it just works, but the lack of control makes me uneasy.
So at some point I want to move it, hosted by myself (with CDN) and under my own domain. Except all the links that people have used to link to my GH pages blog are still out there and I'm not sure how to cleanly redirect visitors to my own hosted version of the blog.
window.location hack? Has anyone else done this (I appreciate my lack of foresight on my behalf is a problem of my own making...)
Or, like you suggested, make a skeleton site out of the current version of the site where each page just has a bit of JS to redirect a visitor.
The nice thing about IPFS is that it has this out of the box. Pages can never die as long as at least one node has them in their cache, even if the original owner went away.
A more comprehensive 404 page (that says that what what you were looking for is not there, and links to the current and the old site), or redirects for the most accessed URLs of the old site in the new one are better approachs in my opinion.
Shouldn't it be 301 rather than 302, though?
If you added a banner at top saying it's "archived content" then that would also solve the issue with people being confused by the redirect that other comments have had.