Do not let your CDN betray you: use subresource integrity (2015)
hacks.mozilla.org
hacks.mozilla.org
I really appreciate the clarity of this post. The author is building up the groundwork without skipping steps that may be obvious to many readers. I of course knew the purpose of a hash before reading the article, but some people don't - and that sentence clearly let those users know why the hash matters without making it less readable for knowledgeable readers.
Writing clarity matters.
It's absolutely important in the security of hashes.
It is very vital in case of SRI indeed, as because SRI also intended to shield from potential MITM somewhere in CDN stack. But SRI is useful in other areas too. For example for really bare version handling and for handling (somewhat) gracefully corrupted cdn responses (including errors, empty responses etc).
For that, avalanche effect is not all that necessary and CRC32 could do an ok job too.
It's not really as useful if you are serving your static assets from the same place as the HTML (and you always use HTTPS) but if you load your js/css on another server SRI can still provide some protection.
[0] https://github.com/waysact/webpack-subresource-integrity
[1] https://github.com/waysact/webpack-subresource-integrity#pro...
Say I have jQuery previously loaded a page that included jQuery from CDNJS and now I'm in China and another site tries to load jQuery from Google's CDN.
Currently that request would get blocked by the great firewall. But since the browser should know that this file matches one it has seen (and cached) before it should be able to just serve the cached file.
This could also save a network request even if I'm linking to a self-hosted file on my own servers if I include the hash.
There is Cache-Control: immutable
https://hacks.mozilla.org/2017/01/using-immutable-caching-to... https://bitsup.blogspot.de/2016/05/cache-control-immutable.h...
As a caveat, there's info leaking here depending on whether the cache hits/misses, so this would need to be opt-in from the cache source, e.g. You set up subresource integrity and also say "allow other domains to load this resource from the cache."
Call it subresource sharing?
Edit: opt-in would need to be on both "sides" (sharer and sharee).
If the site doesn't want to leak that information, it doesn't participate in cache sharing. Since sharing is opt-in, sites won't unknowingly leak this information.
Edit: whoops. I see what you mean. I missed an edit while modifying an earlier draft and left the opt-in only on one side, the sharer.
[1] https://en.wikipedia.org/wiki/Web_Proxy_Auto-Discovery_Proto...
Of course, there are attacks that already work using cache timing, but that isn't a good thing.
Not that many people use CSP, but that's the excuse I've heard for not allowing cross-origin caching.
The problem is that www.victim.example/evil.js doesn't exist, and never did, but your browser won't know that if it's in the cache -- this gives you a way of faking files existing on other servers at the URL of your choice, and as long as they're in the cache you'll get away with it.
and from [2]:
0. evil.example hosts evil.js, <script src=evil.js integrity=foo>.
1. you visit evil.example and the browser stores evil.js with the cache key "foo".
2. you visit victim.example which has an XSS vulnerability, but victim.example thinks it is safe because it uses Content Security Policy and does not allow inline scripts or scripts form evil domains.
3. the XSS attack is loading <script src=www.victim.example/evil.js hash=foo>
4. the browser detects that "foo" is a known hash key and loads the evil.js from cache. Thinking that the file is hosted on victim.example - when the file is in fact not even present.
5. the evil.js script executes in the context of victim.example, even though they use a Content Security Policy to prevent XSS from being exploitable.
[1] https://news.ycombinator.com/item?id=10310594 [2] https://news.ycombinator.com/item?id=10311555 [3] https://news.ycombinator.com/item?id=10312333
(parts first posted here: https://news.ycombinator.com/item?id=13493407#13495482)
[5] https://news.ycombinator.com/item?id=13495482 [6] https://github.com/w3c/webappsec-subresource-integrity/issue...
You create a salt, send it to server.
The server adds the salt to the file, hashes it, then sends the hash back. You check this against the cached data.
When the hash doesn't match, you know either the cached version, or the server version, is wrong.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/If...
This article sums it up really, really well: https://hillbrad.github.io/sri-addressable-caching/sri-addre...
We blogged about Adblock detectors over here[1] a couple of months ago for BugReplay[2]. A major Adblock detector FAdBlock was using subresource integrity to detect whether it's payload was being blocked, ie when adblocker was blocking it from being loaded.
[1]: https://blog.bugreplay.com/2016/11/fkadblock-how-publishers-... [2]: https://www.bugreplay.com
Your stripe js, scary ad networks js, front-end analytics companies. SRI is really neat and helps protect yourself from these many 3rd parties being pwned.
If jQuery is compromised you'll detect it and download from different location but for stripe there is no fallback.
If they want you to stay up-to-date, they'll provide a piece of PHP/Node that emits the latest URL/SRI tag.
Stripe rolls out a fix for a security issue or other bug in their JS. This breaks your subresource-integrity check. The didn't want you to stop accepting payments, they wanted to fix a vuln.
That hampers the usefulness of using subresource-integrity on 3rd-party resources today (which is what yeldarb suggested). Perhaps in the future the 3rd party would provide a script that emits the URL/SRI, but that isn't today.
the 3rd party would provide a script that emits the URL/SRI
And we're back to square one - we can't trust _that_ script to not get pwnedAnd if you are doing any of those things when Stripe pushes an update, how is it any different that having to update the URL/SRI tag?
You might break payment for $polling_interval if the script is incompatible with Stripe's server, so perhaps you could have retry logic there, to bridge $polling_interval more smoothly.
You could also manually review the new Stripe code this way, by polling only by hand or by not automatically updating the cache and SRI.
Then how do you verify the integrity the integrity for the tag?
And anything important like your financial provider is usually very risk-averse to breaking changes, and should give plenty of notice for an update. And if you're including a hash, you don't care about automatic updates anyways.
I can understand trusting one CDN for performance reasons. But do people really add so many different dependencies on their sites? Should I be doing that instead?
I guess the main audience of CDNs are huge sites that will see immediate benefits. For them SRI is good. But using CDN from day 1 seems to me like a premature optimization.
SRI solves some problems (CDN compromise) in exchange for different set of problems (resource not loaded, what now?). And old saying comes to mind... "I had a problem and decided to use regular expressions... Now I have two problems".
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Co...
Ideally this would be combined with something like Binary Transparency, where the new version has to have appeared in a public log for some time, and with no trusted third parties publishing a "Do not trust version X.Y" warning in another public log, acting as a sort of distributed immune system for the web.
This doesn't make sense to me. Why shouldn't I be able to perform integrity checking on resources from non-CORS domains?
See, e.g., <https://github.com/w3c/webappsec/issues/418> for some broader discussion.
The possible solutions to this punctuation-following-URL problem are that you delimit the URL, contort your sentence so the URL is followed by a space and some other words instead of punctuation, start adding random whitespace after the URL but before the punctuation to avoid the linkifier eating the punctuation, or stop putting URLs in plaintext. I've seen all of these used; the first solution is by far the best.
Oh, and that's all from a Western perspective. If, on the other hand, you're using a language that does not use space-separated words (e.g. a number of East Asian languages), then delimiting becomes even more important, because you can't just guess that the URL ends at the space character; there are no space characters around.
I can't speak to your experience seeing or not seeing this syntax, but as I said it's been part of the URL RFCs for over two decades, is used in other RFCs where URLs can appear (e.g. the Link header syntax), and is reasonably commonly used by people who both put URLs in their email and want to punctuate it properly. I will grant that proper punctuation is out of fashion in certain demographic groups. As is writing plaintext, I guess.
[0] https://developer.mozilla.org/en-US/docs/Web/Security/Mixed_...
The other 2 (privacy and authentication) are very important as well and for many are the main reason TLS is wanted.
Plus it still leaks tons of data in the headers. Request time, cache length, useragent, cookies (maybe, hopefully not), accept-* headers, if modified since leaking the last download time, and possibly a lot more.
If you are just loading jQuery you might not care.
If your HTML goes through a CDN (say, you use the full Cloudflare package), the CDN can of course just remove or modify these integrity attributes, or add new scripts altogether.
Disclaimer: I am the developer of sritest