Security advisory: Breach and Django
djangoproject.com
djangoproject.com
* Be served from a server that uses HTTP-level compression
* Reflect user-input in HTTP response bodies
* Reflect a secret (such as a CSRF token) in HTTP response bodies
is vulnerable, regardless of technology.The mitigation strategies were given in the original paper[1], this announcement is just repeat of what's in there. That said, it's exactly the right thing to do, that's not a knock on Django.
this is the only must-have. Last 2 exist almost in every website (everyone needs a CSRF token, everyone reflects something somewhere)
Many web _apps_ have these things, but many web sites do not.
All three parts are a must-have, even if they are incredibly common. Saying otherwise is misleading.
If the application reflects GET request input in the response (eg `https://domain.tld?q=ASDF` results in some `value="ASDF"` being included somewhere in the response), then it is indeed likely vulnerable. This allows the attacker to simply continually change the value of `ASDF` as they guess and check for some secret on the page.
Of course, if your application is allowing untrusted POSTs to be made, then you will still have to worry about POST requests...
On the other hand, if you're saying that we just take this for granted these days, you have made me very happy. :-)
{% csrf_token %}
...results in an insert like... <input type="hidden" name="csrfmiddlewaretoken"
value="566e4606b2094c7c48e5d04b58236f51">
I suspect that the particular mitigation strategy the BREACH authors' describe as "Randomizing secrets per request" could be implemented by having {% csrf_token %} instead emit: <input type="hidden" name="random_data"
value="91178a84e0bc6e08a2fda853eef2d2c8">
<input type="hidden" name="csrfmiddlewaretoken_xor"
value="e0b594e902c7fe6b1748d13aefaf63aa">
...where the random_data changes every response, the emitted csrfmiddlewaretoken_xor is the real token XORed with the random_data, and upon submission the server will again XOR the two values together to get the real CSRF token.There may be other secrets that need protection in other ways, and maybe this would make any random-source issues more exploitable... but this would seem to protect the CSRF token, in a cheap and minimal way.
UPDATE: Thinking further, though, maybe the attacker can probe for both values at the same time, and thus determine the probability of certain pairs, and thus this only slows the attack? I'd appreciate an expert opinion, as this was the first mitigation that came to mind, and if it's wrong-headed I'd like to bash my intuition into better shape with a clue-hammer.
(I'm guessing also, though, it may be possible to probabilistically probe multiple ranges of the secret at once... in a process that seems vaguely similar to forward-error-correction coding.)
1) The attacker must be on the same network as you, or at least be able to detect how large the compressed and encrypted replies are.
If you are on the same network it seems to be there are far more MITM and whatnot attacks that are more likely to succeed, if you do not use HSTS (or secure DNS if that helps).
2) The attacker must be able to get your browser to rapidly generate many (how many?) requests from your browser to the site. It takes "30 seconds" they claim, but is that at a rate 100 requests per second?
3) Each request must carry something that will be reflected by the body of that particular page when it's rendered. I suppose it could be an error message or search string that's echoed.
It seems to me that unless you generate a CSRF token unconditionally on every page, the subset of pages that both reflect something with no protection (e.g. search results) and have a protected form (e.g. change my email address to XYZ) might be small.
4) The secret that can be extracted is what's in the reply body and not the headers -- headers are not compressed, since the TLS compression is now universally disabled post-CRIME.
Personally I use Referer header checking as well. IME all the browsers of my users do send them. So if you extract the CSRF token, it's useless by itself unless you also can make the browser send the right Referer header (and AFAIK, all the holes such as Flash have been plugged).
Other than that -- it seems that if you are normally generating e.g. a 32 byte CSRF key, you could interleave it with 32 bytes of good randomness per request?
Which would be pretty easily averaged out with a few more requests. It makes the attack a little harder, but not substantially so.
I'm not sure how adding 32 bytes of "good randomness" would help.. because the size might be very similar since the randomness might not get properly reduced. And thus, the slight variation in size will still be very relevant.
However, adding between 1 and 32 bytes of randomness might be a pretty good counter! I.e. If you request the page with the "guess letter A", and then request again the same page with the same "guess letter A" and you get +/- 32bytes of different encrypted stuff, it's very hard to assume something was better compressed.
The cool thing about that is that it's fairly trivial to do with most implementation of CSRF. Thoughts?
I was just thinking of hijacking DNS locally (i.e. when browser asks for yourbank.com, send them your IP), and making yourbank.com then redirect to yourbank.myfreehost.com -- which could have a legitimate SSL certificate and a copy of all the branding.
That seems more likely to succeed to me. I guess you could try BREACH if you have some high value target you know is using a very specific website. Anything bitcoin related: either the users themselves or admins. Robbing bitcoins is like robbing banks in the Wild West.
do-not-use-referrer-as-csrf-protection.com
I took the 30 seconds to be the amount of compute time required after they had their samples. Otherwise it is a meaningless number, they could say it takes 0.5 seconds at a rate of 6k requests a second, or 3 thousand years at the rate of 1 request per year.
IME that is not true. Certain corporate/government machines have Referer headers turned off for some strange "security" reason. Last time I ran into this when working with Architect of the Capitol (http://www.aoc.gov/), when they couldn't log into a management panel because Django's CSRF protection checks the Referer header.
EDIT: The following workarounds should be very simple to implement and seem like more viable alternatives for production?
Length hiding (by adding random amount of bytes
to the responses)
Rate-limiting the requests
Mitigations 6 and 7 taken from http://breachattack.com/Disabling compression can break some apps. Especially when they rely on huge compression ratios for text (5-10 times ratio is common for with much json for example). So that is not an app agnostic work around. For example, a 100k of json request, can turn into a 1MB json request. The more data required to send, the more chance of error - especially on 3g/2g networks.
For many high end projects, just disabling compression without regard to testing or having an idea of what the application is doing would get you fired or taken to court.
Not only would this break apps, but it would also lose business in that there is evidence from Amazon and others that every 100ms extra latency can cost 1% in sales.
From SPDY whitepaper: "45 - 1142 ms in page load time simply due to header compression". Remember that headers use the upload part of the link... which means too many headers and you can saturate the upload, therefore making the whole internet connection stall for everyone using it. Common upload limits are only 5-10K/second, so excessive headers combined with many requests can easily DOS many internet connections.
I spend a lot of time optimising websites for these reasons, and disabling compression could add 20 seconds of load time for a good percentage of users.
So, for many apps, turning off compression is no solution at all. You might as well just disconnect your app from the internet - that will also give you a secure and broken app.
A proper risk, and impact analysis should be done first. Too often quick hot fixes to security issues just break things or even make things less secure.
import uuid
csrf = '42be455e20e64d7294eee8d1806d14a9'
p = uuid.uuid4().hex # random response-specific pad
xord = "%2x" % (int(csrf, 16) ^ int(p,16))
request_token = "%s%s" % (p, xord)
print "<input type='hidden' name='token' value='%s'>" % request_token
v_unxord = "%2x" % (int(request_token[len(request_token)/2:], 16) ^ int(request_token[:len(request_token)/2], 16))
if ( v_unxord == csrf ): print "yay, valid CSRF" # constant_time_cmpAnd I guess we can tweak gzip Huffman tables so user inputs were poorly compressed compared to rest of the page content.
For example, when using nginx and with gzip off globally, you can do :
location /static/ {
gzip on;
...
}A very salient piece of cautionary advice. Disable gzip to protect prod. Figure out what to do to allow compression off prod, and engage the devs of your stack/framework to do this correctly.
http://blog.meldium.com/home/2013/8/2/running-rails-defend-y...
The two protective measures are masking the Rails CSRF token and appending a HTML comment to every HTML doc to slow down plaintext recovery. How easy is this to include in a Django plugin?
is there any alternatives ? would like to know what Cloud Flare would do as their CDN is based on compressed nginx responses.
The compressed content of any part of a page very much depends on what came before it. Altering the content to include a script comment block full of random text and various common HTML and JavaScript elements (Markov chains anyone?) would definitely change how a page is compressed.
If the compressed length of the replies varies significantly with every request - even if the request content is identical - attacks like this can no longer reveal hidden information.
Edit:
You could improve this significantly by including false positive matches as well. If your HTML content has: csrf="45a7..." in it, you could hash that content into enough material to generate 19 or so identical looking code blocks embedded in a script comment. You've now provided a 95% chance they attack the wrong one / increased the number of attacks they'll need to try by 20x.
This method (minus the above part) would actually be cacheable by smart CDNs like Cloudflare.
Edit: I think I can see a scenario where a third-party website does these requests via an <iframe> or an <img>. I'm not sure there's a way to do POST quite as easily.
I wrote about this a little while back. Comments are here: https://news.ycombinator.com/item?id=5971464
> DEFLATE [2] (the basis for gzip) takes advantage of repeated strings to shrink the compressed payload, an attacker can use the the reflected URL parameter to guess the secret one character at a time.
By encrypting the CSRF token (or any other "secret" data you want to roundtrip from server to client and back) with a random IV per request this wouldn't work. The value sent by the client would not be the same as the new token generated by the server (since each has a random IV). Even though the decrypted value of each token is the same, the values presented to the client in the response body are each different and not predictable (to the client).
[1]: http://breachattack.com/resources/BREACH%20-%20SSL,%20gone%2...
> We offer several tactics for mitigating the attack.
>
> * Randomizing secrets per request
http://breachattack.com/#mitigationsChunked Transfer encoding is basically padding that a server can easily control, without having to change content or behavior of a backend application. A web server could easily insert an order of magnitude more chunks, and randomly place them in the response stream.
This area of things isn't my strong suite, but assuming that this is analogous to just adding random data to the response, I believe that simply adding random data to the response can be worked around by doing more requests as using statistics to factor out the noise introduced.
If my understanding is wrong then excuse me :)
The difference is that it is extremely easy to add at an http proxy or load balancer level, and could potentially turn 30 seconds into hours;
I would love a way to figure out the math on how many bytes of random response length changes the number of requests needed?
But if random can be statistically removed, then they shouldn't add a random amount. Maybe just track the max size of the returned response and always add enough to reach that max size. Therefore the lengths of all the pages will always be the same. This is still better than turning off compression completely. A typical max for a detail page in an app might just be the size of the page plus 256 bytes per app output field.
http://tools.ietf.org/html/rfc2616#section-3.6.1
So you could make all http responses round into 128 byte chunks, by appending 1 to 128 bytes at the end of every response.
Effectively it gives you length hiding at an http layer; Still attackable.
Is there any general way of preventing this kind of attacks? Inserting random data could work, but it's distribution would have to be exactly right for the attack to be impossible over longer periods of time. For the BREACH case, we could solve it by not compressing user input, but what about the VOIP case?
Also, why does the site http://breachattack.com/ says that "Randomizing secrets per request" is less effective than disabling compression?
Disabling compression is a 100%-effective countermeasure for compression oracle attacks.
> Also, why does the site http://breachattack.com/ says that "Randomizing secrets per request" is less effective than disabling compression?
Putting random data in the server response will only slow down the attack. With enough requests, the noise from that random data will wash out.
Disabling compression will stop the attack cold. The whole thing is predicated on analyzing the size of the compressed text. No compression, no compression oracle.
I suppose, given infinite time, you could send the same request over & over and map the variance of content lengths, and get an idea of what the actual content length was before random padding? But the compression seems to throw that off even more AFAIK - because the data we pad with is random, it could very well accidentally compress well because of the rest of the data in the response, further throwing off any guesses.
Edit: From the pdf on breachattack.com:
While this measure does make the attack take longer, it does so only slightly.
The countermeasure requires the attacker to issue more requests, and measure the
sizes of more responses, but not enough to make the attack infeasible. By repeating
requests and averaging the sizes of the corresponding responses, the attacker can
quickly learn the true length of the cipher text. This essentially boils down to the
fact that the standard error of the mean in this case is inversely proportional to p
N, where N is the number of repeat requests the attacker makes for each guess.Attacker gets secret wrong, page is size: original page + (zero to fifty) + length of incorrect secret Attacker gets secret right, page is size: original page + (zero to fifty)
With a sufficiently high number of observations the attack with the right secret has a mean value that is lower than the attack with the incorrect secret.
At least that's what it appears to me. I could be wrong.
Yes, that's what I mentioned in my post. However, the way I understand "Randomizing secrets per request", it means sending new, random secrets with every request (i.e. generating a new CSRF token for every request).
It's not Django, but us over at Rails have been discussing various parts of BREACH and how we'll handle it: https://github.com/rails/rails/pull/11729
The important two comments are here: https://github.com/rails/rails/pull/11729/#issuecomment-2206... and https://github.com/rails/rails/pull/11729/#issuecomment-2208...
> Let's let this stew for a while with security researchers doing their
> analysis on various approaches and wait and see what the security community
> as a whole recommends.
>
> My only concern to rushing out a release is that we do something equally
> dumb and end up creating a different problem for our users.
>
> We can roll out fixes as it becomes clear what the consensus is as to the
> best solution for a generalised framework like Rails.
As http://breachattack.com/ says: * Be served from a server that uses HTTP-level compression
* Reflect user-input in HTTP response bodies
* Reflect a secret (such as a CSRF token) in HTTP response bodies
These things are easy to tell about your application, but are much harder for frameworks to detect generally, which is why projects like Django and Rails will take some time to evaluate exactly how to best handle this at the framework level.The big bold text ("BREACH may be used to compromise Django's CSRF protection") is a strong warning of the threat (becoming vulnerable to XSS). They list two steps that they recommend taking; disable the gzip middleware in your settings.py, and disable gzip for responses from your web server.
In Django, this means that attackers could recover the CSRF token that's used to prevent cross site requests. This means anyone between you and a client could later have that client automatically make authenticated requests to your app, simply by visiting a site they control, without the knowledge of the user.
To protect yourself, the Django team recommends turning off compression either at the TLS level and at the HTTP level.
> While CRIME was mitigated by disabling TLS/SPDY compression (and by modifying
> gzip to allow for explicit separation of compression contexts in SPDY),
> BREACH attacks HTTP responses. These are compressed using the common HTTP
> compression, which is much more common than TLS-level compression. This
> allows essentially the same attack demonstrated by Duong and Rizzo, but
> without relying on TLS-level compression (as they anticipated).
and > It is important to note that the attack is agnostic to the version of
> TLS/SSL, and _does not require TLS-layer compression._
breachattack.com1) Append a bit to the message before it is compressed and encrypted. 2) See the size of the final message.
So I start by appending the string "4179174b19e0cdc91bf4" to your plaintext message. I see the final encrypted message size is 500 bytes.
Then, I redo the experiment, but this time, I append the string "cschmidt@example.com" to the message. The final encrypted message size is now 480 bytes. The string I injected was the same size, but the compression worked better this time, and I can guess it's because the string I picked is redundant with something in your plaintext.
Mix in a bunch of complicated math and a bit of javascript, and you've got an exploit.
This threat isn't specific to Django: it's being billed as a TLS attack, but any encryption system that uses compression the same way is vulnerable.
It seems like the attack has the following requirements:
1. You want a secret that appears in the response body, like a
CSRF token.
2. The web server always responds with the exact same response
for a request.
3. The response body contains data that you send to the server,
e.g. url params.
4. The attacker has access to an environment where he can send requests
under your browser session (otherwise, the user would be
unauthenticated and there would be no secrets to steal).
Given (4.), how is this a real concern? If I, an attacker, am able to make 3000+ requests while logged in under your session and modify the request character by character pre-encryption, doesn't it logically follow that I have your cookies anyway?Maybe you buy some targetted ads served in an iframe. Maybe you send the user an email where his email server either always shows images, or you trick the user in clicking 'display images' with promise of kittens.
You won't be able to see the results directly, but if you can observe how long the encrypted responses will be, you'll know whether your reflected input could make use of the compression dictionary (meaning your reflected input matches the secret) or not.
I wonder if there is any way to even do this without the passive network snooping -- like some kind of internal browser stats API call that tells you # of HTTPS bytes transferred. It could be innocent enough so it's not protected.
If the system has a decent random-implementation there is no secret involved, just a (pseudo)random string -- essentially a csrf cookie is given the client on one request, and compared on the next request(s).
Is there any reason one couldn't simply use the rotate_token()-function on every (n) request(s)?
For instance, if you had a search field, the contents of what users puts in that search field will not be compromised. However, if you include a csrf token with the search field form, that can be compromised since it will be there every time the attacker gets the victim to make a request.
How hard did you try? Django uses the industry standard security@ address for reporting security issues.
A quick googling results in this page pretty easily: https://docs.djangoproject.com/en/1.5/internals/security/
EDIT: I described the link as 'first' in the Google results, but that was because Google was being helpful and promoting a page I've visited a lot before... In reality, it's a few links down.
> Report potential security issues in Django via private email to
> security@djangoproject.com, and not via Django's Trac instance or the
> django-developers mailing list>Report potential security issues in Django via private email to security@djangoproject.com, and not via Django's Trac instance or the django-developers mailing list
There's also https://github.com/mozilla/django-session-csrf, an alternate CSRF implementation by Mozilla that does use session-linked CSRF tokens. So if you insist on "tokens must be session-linked", you can use that instead.
https://github.com/mozilla/django-session-csrf seems ok, should be default
Perhaps you should do a little research before proclaiming things insecure?
if it's not enough:
some websites from http://www.djangosites.org/ are vulnerable > django has a problem
>Your site is on a subdomain with other sites that are not under your control, so cookies could come from anywhere.
it should be default, for sure.
To the first point, we believe that Django's CSRF protection is as strong as session-linked CSRF protection, and adds CSRF protection to anonymous users (users without a session as well). In other words, it's a design decision, one that we believe doesn't compromise CSRF protection. If you believe otherwise, please get in touch (see above).
2) I checked again for instance https://bitbucket.org/ - edit csrftoken cookie to any, 123123 for example. Reload the page and if site keeps working - Cookie Forcing with MITM will do the same thing using http: injector Set-Cookie
not only MITM, subdomains can do precisely same thing. Either bitbucked uses old django or django is vulnerable to it (which is, well, a severe vulnerability imo)
2) First, you should report this to Bitbucket: https://www.atlassian.com/security. And c'mon, disclosing a possible CSRF vulnerability on a public board is kinda irresponsible. Is responsible disclosure not something you practice?
SecondI don't know what Bitbucket is running, exactly, and exrapolating from Bitbucket to Django is pretty lazy. Frameworks != sites. Once again, we've spent quite of bit of time validating the design and implementation of Django's CSRF protection, and we believe it works. If you find proof otherwise, can you please send it to security@djangoproject.com, and not post it to Hacker News?
>Not all sites use sessions (some for performance reasons, others for privacy reasons);
what kind of site doesn't use sessions? To track a user you need a cookie right?
2) Frameworks != sites. As I used to think, only framework is responsible for CSRF protection, hence I extrapolated. I sent it to security@ as soon as I found this email. I am trying to not proclaim anything but some websites from http://www.djangosites.org/ are vulnerable.
I guess there's the chance that you could do CSRF because you've essentially "set" their CSRF token?
Exactly, cookie forcing/tossing = "set" their CSRF token
1. User is browsing an HTTPS Django site and a HTTP site on WiFi.
2. Hacker is MITMing a connection (on WiFi), cannot decrypt SSL.
3. Changes user's CSRF token for Django via cookie-forcing.
4. Hacker can now use user's other HTTP session, and inject JS (or whatever) so their browser sends stuff to the HTTPS Django site, fully knowing their CSRF token (because we set it) and thus can forge requests easily.
Does Django mitigate something like that already? I think should be pretty easy to mitigate by using Django's signing to sign the session ID or something into the CSRF token?
Cookies are broken (i write about it on my blog like, daily). The essential idea of Forcing is injecting cookies into HTTPS space from HTTP.
http://scarybeastsecurity.blogspot.com.es/2008/11/cookie-for...
Is that possible under HSTS?
> unless it's a subdomain in which you have the more serious issue of session theft or session fixation
elaborate plz? From what I know, subdomain can only do same attack (cookie tossing). Are you talking about phishing?
In any case that you can edit the CSRF token you already can execute a much stronger attack (MITM, XSS, etc). If you have a way to set your own arbitrary cookie that doesn't require a much work attack that already includes the ability to do arbitrary requests without them needing to be cross origin then I heartily suggest you report it.
MITM is not a problem for HTTPS, but it can force cookies. Subdomain can do same thing. These preconditions are common
A CSRF request from a plaintext subdomain would not include a header and would fail Django's CSRF for lacking a referrer header (Strict referrer checking only available under HTTPS). Further more even a subdomain under TLS would fail to have the same origin.
A subdomain can set cookies this is true. This requires a XSS on the subdomain or allowing plaintext responses on subdomains. If you do not ensure both of these then besides being able to set (or in the case of XSS, read) the CSRF token you can also fixate the session, steal the session, preform a DoS using the size of the cookie, etc.
The solution is forced TLS with HSTS and includeSubdomains.
that's true. Few questions: 1) does django provide HSTS header by default? didn't find it in codebase 2) does django has includeSubdomains? bitbucket doesn't have it 3) it is still vulnerable to subdomain-XSS tossing.
taking into account all possible ways to replace the token I think it's django's duty to couple it with session (whatever it is in django) by default. What do you think, should it stay semi manual?
2) The above addon does have it. I have no idea what bitbucket does.
3) Not sure I understand what this means, you mean if there is a valid XSS on a subdomain?
Decoupling CSRF and Sessions has been a requirement for us for awhile now. We've spoken a little bit about optionally coupling them where if you have the session framework enabled it will couple them but if you don't it falls back on the current method.