301 redirects: a dangerous one way street (2012)
jacquesmattheij.com
jacquesmattheij.com
You often can't use 302 because all your external links no longer work SEO magic for you with a 302. Google only transfers link juice with 301 [1].
If you make a mistake and misconfigure your server, you're toast.
If a disgruntled employee 301 redirects your domain, you're toast.
If a service provider misconfigures your domain, you're toast.
If a hacker (from a competitor) 301 redirects your domain, you're toast.
If you buy a domain that had a 301 on it, it's worthless.
If you buy a domain that had 301s on it that point to phishing sites, you're in trouble.
I always add cache headers to 301 redirects I use to at least prevent me from shooting myself with an arrow in my knee.
UPDATE: [1] Google seems to have changed this recently. It also no longer considers http/https different pages as it did in the past with the same content https://www.searchenginejournal.com/google-confirms-no-loss-...
http://mark.koli.ch/set-cache-control-and-expires-headers-on...
Reminds me of the guy that "got hold" of Google's domain for a few minutes. What if something similar happened and someone were able to make this redirect? Millions would be affected.
It seems like browsers interpreting permanent as forever is some kind of a bug. Even if that's literally what it says, that's not what anybody wants. What great evil is being prevented by not having it expire and be refreshed after 45 days?
Also, "This response is cacheable unless indicated otherwise," says RFC 2616.
Working as designed, IMNSHO. Perhaps not working as intended, but alas, that's a case of ¬RTFM.
Much like webpages that say "404 not found" with a "200 OK" header.
But yeah, I agree- (though I've missed them I'm sure) I pretty voraciously make sure my HTTP status matches my intended response- and sometimes that means you've got to write the code in the controller that tells the request to specifically return a certain status.
/quote RFC2616 (So, the status line is an entity of its own, which is followed by headers.)
- I definitely see bad design in interpreting the RFC as "if no caching metadata, then cache forever;" this violates Principle of Least Astonishment.
- Also, "301 and 302 has completely different meaning to a crawler" seems illogical.
- Perhaps the RFC should have also specified something like "in absence of caching headers, don't assume 'cache forever', that is a long time"; it does not try too much to prevent implementors from shooting themselves (and users) in the feet.
In defense of the browsers - the RFC does say, in essence, "rewrite the 301'd URL and be done with it," so the browser does not even need to deal with the redirect any more.
At the least, a redirect should not be permanent across name server changes (when the domain changes hands). Unfortunately, this would effectively be adding state to HTTP.
I did some testing in 2009, think around 2012 and 2014. Additional to loffilegrepping after some big site URL rewrites.
It's a non issue. No caching headers, the redirect gets cached only for the current browser session. Close it, reopen it, gone, done.
Lets discuss this one based on data. (Which I cant provide right now as Im on a beach on Sri Lanka right now with a FirefoxOS device amd I dont know how to see Http requests on this one, but) Please prove me wrong! based on test, data, not blogposts.
I just tested it in Firefox on a Mac. I restarted Firefox. I even rebooted. Developer Tools > Network tab says "cached". I can't confirm that it is cached forever, but it is not only "for the current browser session".
https://dxr.mozilla.org/mozilla-central/source/netwerk/proto...
The 301, 308 stuff comes from IsPermanentRedirect which is here: https://dxr.mozilla.org/mozilla-central/source/netwerk/proto...
I guess some people think the purpose of 301 is more like that of 410: update references so you don't try to go there again. The difference is that with 301 you additionally instruct the client to not even attempt to go there again in the future.
But the article does raise an interesting point: if I own somedomain.example and set it up with a 301 redirect to myotherdomain.example and enough people visit it that most people will have cached the redirect, doesn't that basically mean I now own it for perpetuity (or until enough people have cleared their cache) even if I don't renew the domain and new requests to the domain are no longer served (by the same IP)?
Or do browsers have some kind of protections against this, at least based on DNS? It's a bit too convoluted for a proper DOS attack (because you need to own the domain long enough and make it popular enough to poison everyone's caches) but a naive implementation seems like it would effectively render domains unusable if someone set up a 301 on them at some point in the past.
It also caused half a day of confusion to understand why some of our web browsers were still failing to connect and others could see the alpha site (because they'd never visited the previous 301 site at that address).
Really? In the simplest case, their entire job is certifying that the holder of the private key is the holder of the domain name[1]. That begs the question, of course: how is it that we trust every single CA to certify every single domain? Why don't we trust the issuer of each domain hierarchy to certify only those domains it's permitted to issue?
The entire XPKI is broken, broken, _broken_.
[1] In the more complex case, of course, they certify that the keyholder is some external entity.
https://en.wikipedia.org/wiki/HTTP_Strict_Transport_Security
Although modern browsers are clever enough to detect a redirect loop and throw an error, they're not clever enough to detect when the redirect loop is caused by a cached 301 response. So they cache the redirect loop as well. Throw in another layer of caching (CloudFlare), and now you've got a bunch of URLs that will be stuck in a redirect loop for a very long time.
The only solution was to append some garbage to every URL, like "?cache=no". Fortunately, the problem only occurred with static content, so nginx happily discarded the querystring and returned fresh content.
It would have been useful to include those headers in the blog post.
Last-Modified: Fri, 19 Feb 2016 12:54:49 +0100
Expires: Sun, 29 May 2016 12:54:49 +0200
Cache-Control: max-age=8640000, must-revalidate
See also: https://www.mnot.net/cache_docs/#CONTROL
Also, 301 might poison proxy-caches, so even if you clear the cache in your browser it might still not work.
Whats more surprising is "13.2.2 Heuristic Expiration"[1]
If you specify (or your framework does) a Last-Modified time WITHOUT Cache-Control the browser is free to make up its OWN cache expiration rules (the item is implicitly cacheable)
But then I do accept this perspective (if it's within your call to take this risk). Just don't 301-redirect it. Let the search engines figure it out for themselves.
If a user has a bookmark to an old resource, then it's a liability for you to try to keep your web of 301s working.
KISS!
Also, start with a 302, and only change the status to 301 once you're confident they are correct.
a. most popular websites nowadays are so bloated that browser caches would throw out many things (including your tiny site that someone checks a few times a week or even longer) a lot sooner compared to how things were about 10 years ago (I presume the default disk cache sizes in browsers have not increased by multiples in this period).
b. more people are browsing through mobile devices that are dumped in a few years and replaced with a new one, new browser, empty cache, etc.
It would be good to get an indication of the potential impact.
edit: > A sane client should then check if the old 301 is still there Check out kijin's comment about this.
--- begin factually incorrect statement ---
To save you a click: Jeb Bush forgot to renew his domain,
and Trump bought it and redirected to
www.donaldjtrump.com
Thankfully, Trump is using a 302, not a 301.
--- end factually incorrect statement ---EDIT: ah I see -- jebbush.com has only ever had that redirect throughout its entire (short) existence. Leaving my comment up though, since it's more about what one could do in that scenario than any particular current event.