Do not let your CDN betray you: Use Subresource Integrity
hacks.mozilla.org
hacks.mozilla.org
So your site could use the cached copy of jquery (for instance) that was originally brought down to serve my site, or vice versa.
I don't see why it would be necessary to provide a name and a size. If the hash is the same, we can be pretty sure that's the same file.
If safety were an argument you'd add an extra, unrelated hash function. E.g. even md5 is likely much harder to break if you also have the CRC32, even though CRC is a thoroughly insecure hash (and of course, you wouldn't use an insecure hash, now would you?)
Generating two files where their hashes collide is extremely difficult. Generating two files where their hashes collide at the same size is near impossible, even after you break the hash function itself (e.g. with MD5 it requires much more compute power to generate two files with matching hashes and sizes than just hashes alone, since you're effectively looking for a subset of all collisions).
I'm really glad that the SHA-2 family of hash functions has held up thus far.
(Resource timing is disallowed for cross-domain resources for exactly this reason).
I made a site where I 'guessed' if a user was an employee of a certain company by embedding an <img> from an employee portal (from the login screen) and timing how long it took for onLoad to trigger. Times under a certain threshold indicated that the image was probably cached and they work there.
Of course, once you load it once it's cached, so you need to persist the first result... but you get the drift.
Sharing resources across pages based on has values would render this restriction moot.
From http://www.w3.org/TR/resource-timing/:
"Statistical fingerprinting is a privacy concern where a malicious web site may determine whether a user has visited a third-party web site by measuring the timing of cache hits and misses of resources in the third-party web site. Though the PerformanceResourceTiming interface gives timing information for resources in a document, the cross-origin restrictions prevent making this privacy concern any worse than it is today using the load event on resources to measure timing to determine cache hits and misses."
Unfortunately, this could be used as a cache poisoning attack to bypass Content Security Policy.
See the section about "Content addressable storage" at <https://frederik-braun.com/subresource-integrity.html>.
(If you can come up with a magical solution to this problem, join the W3C web application security group mailing list and send us an email.)
There's a warehouse 'owner' internal to the browser (and not exposed to pages/extension), who 'remembers' the resources the browser has, and the times it took to access them (for commonly-accessed resources). When a page requests the cached resource, the owner 'returns' the resource with a delay of whatever the original access time was, fudged around by some noise.
A weakness of this would be that websites would be able to 'communicate' with each other by engineering response times to your browser, and then checking how long it takes your browser to access that. But this is a different scenario than random websites trying to figure out where else you've been: here the pages need to be in collusion with each other.
The servers might try to use fancy algorithms to try to figure out if you're using cached versions by hitting different distributed servers and figuring out if the resource load time is an outlier. But that's prone to a lot of noise and other issues, and lesser of an issue than the original concern. Right?
This obviously won't help with page load speeds, but will help network load for bot users and servers. One possible issue might be: if you've been in a slow connection previously, all your connections will seem slow even after you are in a faster connection. For that, you can just purge the cache and force the browser to reload the resources.
Edited for formatting.
A site operator would only opt in to the behavior for assets that are not unique to the site.
A second idea would be to wait until several unique domains had requested the asset before turning on the behavior for that asset. (By unique domain I specifically mean the part of the domain that's written in black text in the URL bar, excluding subdomains that are in gray.)
These are two easiest-to-implement solutions I can think of.
How about an additional attribute named "global",
"shared", "public", "use-global-cache", share-with="*",
etc. that the developer can use to opt in to the behavior?
Allowing people to opt-in to a cache poisoning vector seems like a bad idea. A second idea would be to wait until several unique domains
had requested the asset before turning on the behavior for
that asset.
This just raises the bar to a cache poisoning attack from "owns one domain name" to "owns a couple". Some gTLDs are $0.99 per year, or free. (The user would only have to visit a single page, which has a dozen other sites open in invisible iframes)If the hash-cache wasn't there, then good.com would have returned a 404. There's no hash collision because the request is completely elided (which is a large part of the perf attractiveness).
And as for the "collisions are unreasonable to expect people to generate", remember the use case: these are going to be extremely long-lived hashes.
With the cache poisoning, once you find a collision against jQuery 2.1.1 (to beat the example horse), you can continue to use that against all requests for jQuery 2.1.1. And we know how wide-applicable targets of cryptographic opportunity typically fair against adversaries with substantial brute-force processing resources...
The impact of SHA-2 failing would be far, far, larger than poisoning jQuery.
What would the attacker here be doing, and how? I read the piece on your site, and it's not clear to me what the attacker would be updating, and what effect it would have. Can you explain?
1. you visit evil.com and the browser stores evil.js with the cache key "foo".
2. you visit victim.com which has an XSS vulnerability, but victim.com thinks it is safe because it uses Content Security Policy and does not allow inline scripts or scripts form evil domains.
3. the XSS attack is loading <script src=www.victim.com/evil.js hash=foo>
4. the browser detects that "foo" is a known hash key and loads the evil.js from cache. Thinking that the file is hosted on victim.com - when the file is in fact not even present.
5. the evil.js script executes in the context of victim.com, even though they use a Content Security Policy to prevent XSS from being exploitable.
Most attackers don't have that capacity, though; XSS is usually done by tricking the page into running your own JS code (for example, by finding a publicly editable text area which doesn't properly escape HTML). Those restrictions wouldn't stop this attack.
EDIT: Although I guess the current URL based cache may already dedupe, in which case my solution would be roughly equivalent to just turning off hash-based caching for domains with CSP.
The original poster said to also use the size. If you include that, my understanding is that crafting a hash collision is into the realms of impossibility.
Am I wrong?
===========
Turns out I was completely misunderstanding. I now do. Thanks!
3. the XSS attack is loading <script src=www.victim.com/evil.js hash=foo size=123>
> If you include that, my understanding is that crafting a hash collision is into the realms of impossibility.
Crafting a hash collision is already in the realm of impossibility. (They're using cryptographic hashes: if you can make SHA256 collide, we have bigger problems.) The attack here isn't that you're getting the wrong file, it's that you're getting a file the webserver does not have, at all. Step 4 is where we go wrong: we load the file from cache, while we should instead request it from the server, which will 404 the request because it does not have the file.
(And it's JS: even if size did matter, you can just add spaces to the end…)
victim.com should be protected because it's content security policy tells the browser not to run scripts from evil.com, but the browser thinks that evil.js came from victim.com, even though victim.com doesn't host evil.js and the browsers cache got evil.js from evil.com.
Edit: But if we only make the global cache work on the same domain, this problem should disapear completely (it's obviously not as powerfull then but still a massive improvement to the current system)
How so? Browsers have had caches for consistent URIs for some time now.
What do you think are the downsides to something like that?
Given something like this:
<script src="https://code.jquery.com/....." integrity="sha384-R4/....." shared>
(where 'shared' invokes the caching mechanism)The browser sends a request for it with If-None-Match: "sha384-R4/....." header set.
I think this solves 99% of the problem:
If the integrity tag doesn't match the ETag of the resource, the server interprets it as out-of-date cache and responds with content of that resource. If the integrity tag matches the ETag of the resource, it will respond with '304 Not Modified'.
And that's the remaining 1% of the attack surface: basically the attacker wins iff the site can be tricked into serving a resource with the same ETag as the hash of his payload. We don't need to worry about collisions: even if someone uses ETags that match the form of subresource integrity tags without intending to, the attacker would still need to generate a collision, which is just as hard as finding a collision with any other hash. But if there are servers out there that will serve files with externally-set ETags then they'd be exploitable.
Just do a request and hope that the server returns you a 304 Not-Modified.
This should prevent the issue, right?
> Content injection (XSS)
If we assume XSS, an attacker could simply inject whatever they want. The cache isn't needed. This still wouldn't poison any legitimate cache keys.
> The client still has to find out if the server really hosts this file.
So use (URL, hash) as the key in the permanent cache. This removes most of the bandwidth, and using a CDN allows for one GET per file across many sites.
So what exactly is the attack? I'm really not seeing how someone could attack a permanent cache without first breaking the hashing functions that we already have to trust.
edit: after reading https://news.ycombinator.com/item?id=10311555
This would work in the cases where we allow XSS (which is already a compromised scenario). Simply adding the URL (or maybe even just the hostname) prevents this entirely, and we still get almost all of the benefits for local resources, and we get all of the benefits when using a CDN.
edit2:
There are two issues being discussed. 1) Is the file we loaded form a (possibly 3rd party) site correct? 2) Did we ask for the correct file(s).
Cache poisoning is when you can fool #1, while XSS attacks manipulate #2.
But there are also parts important to interpreting the file that aren't part of the hash, like the mime type. I think this problem is a lot more complicated than you're saying.
Hence, the attack here is making the victim load that keyed script on a different page, then redirecting them to an XSS hole that links to that script as 'hosted' by a whitelisted domain. Since it seems to be on a whitelisted domain and match the original script's hash, it will execute on the page, which is not ordinarily possible on a page which is running CSP.
I hope this encourages you to not immediately assume that large groups of people working on technically complicated problems are stupid in the future.
The scenario in question has Protected Site A vulnerable to an XSS attack, but protected from it due to their CSP not allowing scripts from foreign domains (only trusting scripts from `trusted.example.com`). This is what CSP is for: it's not for what you expect to serve, it's for protecting against what you don't expect to serve.
In the theoretical attack content-addressable scripting could open up, the user visits Malicious Site B, which loads a malicious script with the hash `abad1dea`. The owners of Malicious Site B use their XSS attack to insert the (simplified) HTML `<script src="https://trusted.example.com/payload.js" hash="abad1dea">`. If Malicious Site B tried to insert a direct link to their payload at `malicious.example.com/payload.js`, it would be blocked due to the site's CSP - however, if the site trusted the fact that it's seen `abad1dea` from `malicious.example.com` as evidence that it could get the script from `trusted.example.com`, this would open up a vector allowing Malicious Site B to run the `abad1dea` payload in a way that would not be blocked by the CSP. This is why the UA still has to make the request, even though it already has the content.
With the behavior that's been specced, a request will be made to `trusted.example.com` which will either 404 or give a different script, causing the XSS attack to be blocked by the page's CSP.
An attacker can inject whatever they want, but they can't run whatever they want. That's the purpose of a Content Security Policy: the problem isn't the content of the script being run, it's the context in which that script is being considered.
Because different scripts are given different permissions (eg. access to cookies) based on their domain of origin, the existence of said content must be verified to be true in the context in which it asserts its presence.
It's not a cache-poisoning attack so much as it is a cache-use attack, but it is a legitimate attack.
cache = {
b0af301e782bf5e2a8ccce919b6ca3b70aa771db:
{domains: ['evil.com','airbnb.com'], content: '...'},
35d778783c4155c20360d269c9dd000fdcd39548:
{domains: ['javajosh.com'], content:'...'}
}
You go to secure.com, but a malicious user has put the b0af301 script in your path. CSP's white list for secure.com is [secure.com, javajosh.com]. The browser dereferences the hash, checks against the associated domains, and rejects if a whitelisted domain isn't in that list. Your browser running secure.com would reject the b0af301 script.(Something I personally would like would be for for orgs like EFF.org to post known-good hashes, so I can always add the EFF hashes to my site's CSP whitelist, and have a warm-and-fuzzy feeling.)
I think this would also allow the server to "pre-validate" with HTTP2 push.
1. Add an If-Hash-Mismatch header so you don't need to transfer the body
2. Add a list of hashes to accept to the content security policy headers
3. Add a list of public keys to accept to the content security policy, and allow the content if it's signed by one of those (this requires some standard way of signing things, maybe PGP/MIME or a dedicated HTTP header)
4. Only allow this from <script> and <style> tags that are in the <head>, or that are at "end" of <body> (meaning there are no tags other than <script> or <style> afterwards), or resources referenced from CSS and JavaScript files loaded that way.
EDIT: 5. Add an ECC public key (Curve25519?) to the content security policy, and accept hashes where an extra attribute is specified providing an inline signature of the hash with the key
The idea of the last one is that XSS would usually happen in in the middle of the body and not in the head or footer.
That said, you can XSS with inline script, so it seems this only mitigates XSS vulnerability with length limitations on the payload (EDIT: nope, CSP blocks inline script).
But 1. should exist regardless to complement the existing caching options. It shouldn't be sent by default to avoid adding another tracking method, but if the source page specifies a hash, and you have that hash, then If-Hash-Mismatch is perfect.
2. Bingo, winner.
"pretty sure" isn't really good enough.
Allowing this sort of code sharing provides another attack vector. If a malicious site is able to exploit a browser vulnerability that allows it to populate the shared cache, then you've suddenly enabled code injection on any site using subresource integrity.
Source: https://w3c.github.io/webappsec/specs/subresourceintegrity/#...
I'd be surprised if 3rd party providers don't start intentionally adding a random byte on each request (or every hour or something) to make sure that webdevs don't take a dependency on the contents of their files.
If the CDN doesn't support CORS and the browser does support subresource integrity, subresource integrity is ignored (bad, since an attacker can disable CORS before changing the js) or enforced, thus refusing to execute the js (good)?
SRI returns false (i.e. non-matching integrity) for scripts (or stylesheets) that do not enable CORS and are not same-origin [1]. Otherwise, an attacker could just disable CORS to bypass SRI.
flies away
[1] https://w3c.github.io/webappsec/specs/subresourceintegrity/#...
I suspect this is going to become a turbo AdBlock. If the original page doesn't sign the content, block it.
But with SRI, it should be possible to send scripts, css, etc. over plain HTTP right? As long as the landing page is HTTPS, and the hash checks out, is there any reason for browsers to show a warning to users then?
This exposes a general issue that sometimes you want data integrity but not privacy. With https it's all or nothing.
This seems like exactly the thing that they were talking about when they started depreciating http [0]. Does this mean they've changed their mind?
[0] https://blog.mozilla.org/security/2015/04/30/deprecating-non...
Pages which use this should detect subresource integrity fails and report them to both the browser user and a non-CDN logging machine. Subresource integrity should put a stop to CDNs and ISPs inserting ads and spyware, because if even a few major sites use subresource integrity, they'll get caught quickly and will suffer bad publicity.
This encourages using a CDN for only the bulky parts of a site. Put the important pages (entry pages, login pages, credit card acceptance) on a server you control, with your own SSL cert. Put the resources loaded with subresource integrity on a CDN. Now you're not trusting the CDN at all.
Tools for website maintenance will need some improvement. Files need version info in their names; if the content changes, the URL should change, too. Maybe use the hash as part of the URL. Such files can have indefinite cache expiration times; they're immutable.
Presumably it's not for collision avoidance, and it's not like anyone's going to be hitting the maximum message size of SHA256 with anything stored in CDN..
Edit: So it seems that all of the main variants of the SHA-2 family must be supported[1], and the spec supports multiple hashes being presented at once. It's just that SHA-384 seems to be used in all of the examples I've seen so far.
> Conformant user agents MUST support the SHA-256, SHA-384 and SHA-512 cryptographic hash functions for use as part of a request’s integrity metadata, and MAY support additional hash functions.
> When a hash function is determined to be insecure, user agents SHOULD deprecate and eventually remove support for integrity validation using that hash function. User agents MAY check the validity of responses using a digest based on a deprecated function.
[Edit] I think this was the discussion, just over a year ago! - https://news.ycombinator.com/item?id=8359223
EDIT: I just reread my post and there are some wonky ideas mixed in with (what I think are) decent ones. Sorry!
P.S. Blake2s would be a great additonal hash to support.
2. The people who do hand write HTML tags tend to be precisely the type of people who would go out of their way to generate an md5 checksum on the commandline, or write a script to post process their HTML files.
So, an extra request per resource, in other words?
Requiring CORS simply makes security more difficult to achieve.
See https://annevankesteren.nl/2015/02/same-origin-policy and http://w3c.github.io/webappsec/specs/subresourceintegrity/#c...
Fair enough, that's some information leakage, but it's certainly not easy to exploit. Normal cross-origin limitations still apply, so you'd need to get creative to even get the information in the first place, and if you can, it's not clear what this adds over a timing attack.
I'm still a little skeptical such a heavy handed restriction is necessary to maintain the current level of security; but then again - why take the risk?
The only reason why you want to include third party scripts is that you want automatic bug updates etc!?
The main purpose for CDN hosted scripts is allowing them to be cached in your browser, reducing latency via close edge servers and load times across website via caching.
This proposal prevents malicious changes, as well cache poisoning, which was a very scary threat up until this announcement due to attack vectors like the one described in this defcon talk: https://www.youtube.com/watch?v=kLt_uqSCEUA
(PDF slides here: https://defcon.org/images/defcon-20/dc-20-presentations/Alon...)
Hopefully, subresource integrity schemes will eventually allow browsers to fetch from local cache based on the checksum of the file contents, rather than the URI from which the resource is served. :)
(for iframes, I think this would work fine)
[1]: I'm presuming you're not using TLS, because if you were, then TLS would do this and more.
It raises the bar though from blind injection, which I think is quite good.
It would add overhead to verify hashes in the manner that you have mentioned but I think it's worth it.
We reached out to jQuery and code.jquery.com does this for a few months now.
If I have the sha-256 of an exe file, I'm perfectly happy to run any exe file with the same sha256 simply because collisions don't happen. Why is this different for JavaScript?
If an attacker can inject HTML script tags into your website haven't you already lost?
<script ... integrity="sha384-...">
How will the browser confirm that the source HTML requesting script isn't modified by the CDN?Well known CDN services, such as CloudFlare, are known for modifying the HTML -- This is not uncommon.
It might be different for other sites/stacks but on ours we deliver CSP at the whole site level, meaning it is delivered with every response we send.
Script integrity is only sent when that specific script is used, and it means our workflow doesn't have to change to rewrite a HTTP header dynamically with each page (based on which scripts are or aren't on that specific page).
I legitimately have no idea how I would implement CSP with hashes for the scripts on that specific page. It would require me to actually patch the software stack upstream. I do however know exactly how I'd use the integrity field on a script block and could implement it with just raw HTML.
PS - Not to mention that few browsers support level 2: http://caniuse.com/#feat=contentsecuritypolicy2
However, traditionally, the CDN now controls the content of your JS, and could inject whatever they want into it. That's where this proposal comes in…
(Of course, if it isn't in the cache, then you might need as much as a DNS lookup+TCP connect+a TLS handshake to another host… tradeoffs. HTTP/1.x is also limited to n connections to a DNS name at time, so you can parallelize requests by hosting across multiple domains, such as a CDN, but I find this argument less compelling.)
If average page size is 500 Kb and jquery size is 30 Kb gzipped you do not save much by hosting is on a CDN. What you get is more DNS requests, more downtime when that server fails or stalls and give out data about your users.
I think it is easier just to host everything on your own server.
Congratulations Mozilla!!! This is one of the few recent changes I've seen to browsers that fundamentally changes web-page security in a simple and novel way.
SRI does little for trust when your CDN is also proxy-caching your HTML (e.g. Cloudflare).
If a site wants to pin a third party to a specific version, they'll need to copy the file themselves. Though I'm not sure if this can be detected and "fixed" by the script author. I've noticed that Stripe's js file logs a warning if it thinks it's being loaded from another domain.
Incidentally, I always wonder about this for non-HTTPS sites that offer binary downloads and crypto hashes to verify the files. How can you be sure someone isn't MitM'ing you?
CDNs are used to optimize traffic cost but hosting just single JS library there won't save you much.
Calculating hashes is bothersome and requires modification of an app so probably nobody is going to use it.
I never use libraries hosted at free CDNs.
EDIT: I cannot think of a scenario where this feature can be useful.
EDIT 2: And you cannot use this feature for scripts like Google Analytics because they can be modified anytime.
Using this it looks like you could put the hash in your markup and if their JavaScript code changed it wouldn't execute. For some people having that deadman's switch might be better than an always-execute policy on JS that you aren't writing yourself.