Where did all the HTTP referrers go?
smerity.com
smerity.com
A few weeks ago, I ran across a website (can't remember their domain) that refused to show me any content unless I enabled third-party cookies, and even contained a lengthy argument explaining why disabling third-party cookies hurts the web. IIRC the whole argument was 24K bullshit, written by people who feel entitled to keep making money with their outdated business models and who were obviously alarmed because modern browsers were doing sensible things. And now we're seeing a very similar argument, only this time it's about referers. Why do you think it matters whether I came to your site via Google, Reddit, HN or someone else's blog? What makes you think gives you the right to know that?
Webmasters never had the right to know where your visitors were coming from, any more than the owner of a random gas station on the Interstate has the right to know which city his customers are driving from. If SSL is making Referer headers disappear, good riddance. We just closed a privacy hole, 99 more to go. Next in the TODO list: get rid of referers even when the referring website doesn't use SSL, because as the article correctly points out, we've got a bit of inconsistency there.
Why? It actually solves your "next in the TODO list". Most web sites that shouldn't send referrers don't use <meta name="referrer" content="never"> so will be leaking referrers to other web sites. Adding this meta tag will eliminate referrers in both HTTP and HTTPS.
So yes, my own preference is to keep HTTP Referrers, but I also explain how to kill HTTP referrers for webmasters who would like to as well.
Unfortunately I still can't remember the website where I found the "bullshit" argument, and it's not in my history because I probably ended up using a different browser to comply with their no-access-unless-you-accept-3rd-party-cookies policy. But if you understand why people like me want third-party cookies to be disabled by default, I think you'll also understand why I want referers to be disabled by default, too. It's not about user control as @untog suggests, because the user is always ultimately in control when it comes to HTTP headers. Rather, it's about having secure defaults.
I'm not saying that I totally agree with kijin, but I don't think your answer addresses his point.
It's always the user's choice as to whether to send referrers or not, as the referrer is actually added by the user's web browser itself. Extensions exist for just about every major web browser[1][2][...] to modify the behaviour of the HTTP Referrer field. If you don't like the idea of sending referrers, it's entirely within your control to never send a single referrer.
In almost all cases, disabling the referrer entirely won't result in any broken behaviour, primarily as the HTTP Referrer is unreliable and can be spoofed anyway.
[1]: https://chrome.google.com/webstore/detail/referer-control/hn...
[2]: https://addons.mozilla.org/en-US/firefox/addon/refcontrol/
It is, like you said, basically a massive privacy hole. But a very handy one
URL fragments aren't transmitted to the server, so that means you lose referrer analysis for static sites.
As an occasional small-time blogger, I like to know where the conversation happens, where people are interested - it enhances my engagement with my readers. As a reader, I'll gladly grant webmasters that courtesy.
Personally though, I'd rather not block hotlinking. But I understand why some people are against it.
The ability to effectively DOS anyone's site in this way, not to mention presenting their content as your own without technically infringing copyright (or so the Ninth Circuit seem to feel in the US, though courts in other jurisdictions have differed) is a genuine and, for the unfortunate victim, potentially very expensive problem with the current state of the web. Doing this has always been bad netiquette, but these days even Google do it, and indeed used the fact that they were doing it in their defence in one of the aforementioned copyright cases.
http://ascii.textfiles.com/archives/1011
(No goatse images on the page I directly linked to. It does link to them, but the links are marked.)
I host a small, very low traffic website. One day, the bandwidth shoots through the roof and stays high. The reason? One of the images on a page got added to the .sig of someone in a popular forum. Suddenly thousands of people are fetching the image.
The solution was to filter by referrer header, letting the image be seen by visitors to the actual page, but linking from other sites gets blocked. Note that usually it's best to allow requests that have no referrer header at all, otherwise you'll be blocking some legitimate viewers of your site.
End result: bandwidth back down to the usual, tiny levels. It's not that I cared about people copying the images, I just didn't want to foot the bill for the traffic!
> Note that usually it's best to allow requests that have no referrer header at all, otherwise you'll be blocking some legitimate viewers of your site.
It's just another way that (for some reason) my browser leaks information without asking me if I want to.
It gets switched off now, along with cookies, analytics and most other stuff.
I also respect a site owner's desire to track your referer, and to decide the rules for who gets to access their content. They are (often) a business, and knowing where to focus their efforts to make money is important to them.
I don't respect either side feeling they are owed tracking information, or content, if they aren't willing to respect the other agent's rules and preferences.
1. Follow link from https://example.org to http://example.com --- referrer is not sent
2. Follow link from https://example.org to https://example.com --- referrer is sent
I don't understand how the same referrer can be too sensitive to be sent as plaintext, but harmless enough to be passed to a not-necessarily-trusted third party.
1. Follow link from https://example.org to http://example.com --- can be read by a third party if referrer were added
2. Follow link from https://example.org to https://example.com --- cannot be read by a third party so referrer can be added
The assumption is that secure pages are secure for a reason, and that the author of a secure page is linking to other secure pages and has some basis of trust by which the link is provided.
Let me rephrase my question: why the default assumption that example.com is trusted not to misuse referrer information merely because example.org provides a link and the human user follows that link?
I disagree. When you click a link on a page that you retrieved from example.org, one that leads to example.com, there is no communication between you and example.org, nor between example.com and example.org. The communication that takes place is between you (party 1, the initiator of the conversation) and example.com (party 2, the target). The HTTP request mentions example.org, but being a third party, it does not participate in it directly.
The only conversation in which example.org was a party was the one in which you requested the page that contained a link to example.com, which has already finished.
In that light, it seems strange to me that under HTML5 (assuming I understand the article correctly), example.org is given a mechanism to dictate how much information you give to example.com. Should that not be your choice, as the sender of said information?
2) Page from Site A suggests what should be sent in the referrer via the meta referrer
3) User clicks on link from Site A to Site B
4) User's browser requests page from Site B (referrer is set by either user's overriding option or the meta referrer from Site A)
So indeed, at no point does Site A speak to Site B directly. The meta referrer simply asks the user to either send or not send the referrer. If the meta referrer is not present or not supported, it falls back to default HTTP Referrer behaviour.
As the user, you can override this behaviour and force the referrer to do whatever you'd like. This includes refusing to send it, always sending it, or spoofing it. Firefox for example allows you to set network.http.sendRefererHeader and there are various browser extensions for any popular browser that will allow for finer grained referrer control.
If the destination is untrusted, then the source can just anonymise the redirect by sending it through a point that won't reveal the precise source. This is how services like http://anonym.to/en.html work.
An example where the default behaviour may be appropriate is Facebook interacting with third party apps.
Facebook may be happy to pass referrer information across to these third party apps as long as they handle it securely. If the referrer goes across HTTP, it goes across the Internet in plaintext (unsecure). By ensuring it travels over HTTPS, you're at least ensuring a minimal level of security.
(If this is the protocol we're building economies on these days, I feel computer security is going to get much worse before it gets better...)
Let me expand my argument a little.
We have two security domains, associated with the first and second server. The first server has sensitive URLs, the second receives these in the referrer header.
In making the first connection, the client and first server get to make some security policy decisions: the crypto used in securing the connection, the availability of countermeasures against known TLS protocol vulnerabilities, client and server authentication methods, etc. These parameters are complex, and you end up with a connection with some security level given some attacker model. If you think about it really hard, you can come up with a quantitive estimate of what security you might get from the connection in terms of attacker work -- anywhere from zero (trivially broken) to 256-bit security (very good).
Assume our first connection has a 128-bit security level and has countermeasures for TLS vulnerabilities. This is good going. Now we serve over this connection a https link, which the user clicks.
Unfortunately the client and/or second server is poorly configured and we only manage (for the sake of argument) a 40-bit security level and no TLS vulnerability countermeasures. Now we've reduced portions of the first connection to the security level offerred by the second connection. This is really very surprising.
I see what you mean about dropping the security level, but generally SSL is seen as a binary 'good enough/not good enough' choice. I don't know of any browser that gives a graduated measure of a site's security. Either it flags up a warning or it doesn't.
I've found the information in my referrer logs quite interesting and useful, despite the Russian referrer spam, and am sorry to see it going away.
I still feel that removing referrers entirely destroys many useful tools and analytics that we've traditionally been able to use. It removes the core way in which we understand connections across the Internet. By removing referrers, the best we can do is use link graphs, falling back to the original PageRank algorithm where we assume people are random bots that click on one of the links on the page.
Edit: Can't reply to you hnriot due to comment depth limit. My reference to PageRank is as links and backlinks could be used as a poor referrer substitute, though they aren't currently used as it's a lot more work and less accurate. I'm simply saying that, in the event that referrers all disappeared tomorrow, you'd see normal websites trying to estimate where their traffic comes from by using a PageRank inspired algorithm, or more naively by looking at who links where.
I did, however, think that bringing PageRank into the argument was ill advised and probably detrimental to his goal of advocacy through education. It just confused things, introduced another rabbit-hole concept.
If the author thought that mentioning PageRank would lend credibility to the 'link counting' / 'random link clicking bots' foregone conclusion he setup., he was right. It did. So link the text to a footnote referencing PageRank and stay on topic.
2620:0:1000:3509:ed78:b06d:4942:90e1 - - [28/May/2013:17:35:00 +0000] "GET /software/haskell-dbus/ HTTP/1.1" 200 5819 "https://www.google.com/search?client=ubuntu&channel=fs&q=haskell+dbus&ie=utf-8&oe=utf-8" "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:21.0) Gecko/20100101 Firefox/21.0"http://searchengineland.com/google-puts-a-price-on-privacy-9...
They implement this in two ways: (1) If you go directly to google.com and type in your search, the results page uses a # in the url which keeps all the query parameters out of the referrer. (2) Google has used (not sure if they still do/randomly test whether or not to) JavaScript redirects which overwrite the url when a search result is clicked. I'm sure there are other ways for Google to hide the referrer -- plus Google and various browser extensions can turn parts of this on/off however they choose.
It is still possible to wind up with a referrer from a Google search where you can see the search keywords, if for example the search is done using the browser address/search bar, and the JavaScript overwriting result urls is not active (turned off by NoScript, etc.). However, this is not in Google's best business interest (if they can convince people to pay for the info) so I am counting on them trending towards making this the least likely of possible scenarios.
And I'm running my blog on HTTP.
How is this so?
What websites your visitors consume is IMHO not your business.
But if your product is the subject of discussion on another page which links to yours, you may want to be able to join the discussion.
As a site owner, I don't think it's unreasonable to ask my visitors how and where they found out about my website.
As far as analytics its all done with tracking cookies now, which is another issue that needs to be addressed, but Im not going to miss referer should it ever really go away.
That said, I can see why certain people might see it as problematic... if you're browsing http://www.anarchistsbombmakingforum.com and somebody posts a link to http://www.fbi.gov, then maybe you don't want the FBI knowing you were at anarchistsbombmakingforum.com. But, still, barring other privacy problems, the FBI don't know who you are when you visit their site, just that you came from anarchistsbombmakingforum.com.
I have a hard time getting worked out about this though... for one, if you're surfing anarchistsbombmakingforum.com, common sense would dictate that following a link to fbi.gov isn't such a good idea (and a forum that fosters discussions of anything controversial should probably munge links to go through an anonymizer anyway) AND the people for whom this really matter are the people who have a referer blocking plugin installed in their browser.
The problems are numerous. I think I went through 9 revisions of the "naive protocol" (including, at one point, ditching the referer header in favour of HTTPS). Subsequently I realised that my tracking protocol was broken anyhow.
My conclusion is that there's no reliable way to track users visiting multiple websites using the standard features of HTML/JS/HTTP in the face of malicious users or publishers.
You have to fall back on traffic analysis.
I developed a successor technology which works better in many respects (but not all). As it's the subject of a current patent application I can't really go into much detail.